The Experts below are selected from a list of 49827 Experts worldwide ranked by ideXlab platform

Fileno Alleva - One of the best experts on this subject based on the ideXlab platform.

  • the microsoft 2017 conversational Speech Recognition System
    International Conference on Acoustics Speech and Signal Processing, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas Stolcke
    Abstract:

    We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.

  • The Microsoft 2017 Conversational Speech Recognition System
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.

  • an overview of the sphinx ii Speech Recognition System
    Human Language Technology, 1993
    Co-Authors: Xuedong Huang, M.-y. Hwang, Fileno Alleva, Ronald Rosenfeld
    Abstract:

    In the past year at Carnegie Mellon steady progress has been made in the area of acoustic and language modeling. The result has been a dramatic reduction in Speech Recognition errors in the SPHINX-II System. In this paper, we review SPHINX-II and summarize our recent efforts on improved Speech Recognition. Recently SPHINX-II achieved the lowest error rate in the November 1992 DARPA evaluations. For 5000-word, speaker-independent, continuous, Speech Recognition, the error rate was reduced to 5%.

  • the sphinx ii Speech Recognition System an overview
    Computer Speech & Language, 1992
    Co-Authors: Xuedong Huang, Fileno Alleva, Mei Hwang, Ronald Rosenfeld
    Abstract:

    In order for Speech recognizers to deal with increased task perplexity, speaker variation, and environment variation, improved Speech Recognition is critical. Steady progress has been made along these three dimensions at Carnegie Mellon. In this paper, we review the SPHINX-II Speech Recognition System and summarize our recent efforts on improved Speech Recognition.

James Droppo - One of the best experts on this subject based on the ideXlab platform.

  • the microsoft 2017 conversational Speech Recognition System
    International Conference on Acoustics Speech and Signal Processing, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas Stolcke
    Abstract:

    We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.

  • The Microsoft 2017 Conversational Speech Recognition System
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.

  • the microsoft 2016 conversational Speech Recognition System
    arXiv: Computation and Language, 2016
    Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.

  • efficient on line acoustic environment estimation for fcdcn in a continuous Speech Recognition System
    International Conference on Acoustics Speech and Signal Processing, 2001
    Co-Authors: James Droppo, Alejandro Acero, Li Deng
    Abstract:

    There exists a number of cepstral de-noising algorithms which perform quite well when trained and tested under similar acoustic environments, but degrade quickly under mismatched conditions. We present two key results that make these algorithms practical in real noise environments, with the ability to adapt to different acoustic environments over time. First, we show that it is possible to leverage the existing de-noising computations to estimate the acoustic environment on-line and in real time. Second, we show that it is not necessary to collect large amounts of training data in each environment-clean data with artificial mixing is sufficient. When this new method is used as a pre-processing stage to a large vocabulary Speech Recognition System, it can be made robust to a wide variety of acoustic environments. With synthetic training data, we are able to reduce the word error rate by 27%.

George Saon - One of the best experts on this subject based on the ideXlab platform.

  • The IBM 2016 English conversational telephone Speech Recognition System
    Proceedings of the Annual Conference of the International Speech Communication Association INTERSPEECH, 2016
    Co-Authors: George Saon, Tom Sercu, Steven Rennie, Hong Kwang J. Kuo
    Abstract:

    We describe a collection of acoustic and language modeling techniques that lowered the word error rate of our English conversational telephone LVCSR System to a record 6.6% on the Switchboard subset of the Hub5 2000 evaluation testset. On the acoustic side, we use a score fusion of three strong models: recurrent nets with maxout activations, very deep convolutional nets with 3x3 kernels, and bidirectional long short-term memory nets which operate on FMLLR and i-vector features. On the language modeling side, we use an updated model "M" and hierarchical neural network LMs.

  • the ibm 2015 english conversational telephone Speech Recognition System
    arXiv: Computation and Language, 2015
    Co-Authors: George Saon, Steven J Rennie, Michael Picheny
    Abstract:

    We describe the latest improvements to the IBM English conversational telephone Speech Recognition System. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result.

  • the ibm 2015 english conversational telephone Speech Recognition System
    Conference of the International Speech Communication Association, 2015
    Co-Authors: George Saon, Steven J Rennie, Michael Picheny
    Abstract:

    We describe the latest improvements to the IBM English conversational telephone Speech Recognition System. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. Index Terms: recurrent neural networks, convolutional neural networks, conversational Speech Recognition

Andreas Stolcke - One of the best experts on this subject based on the ideXlab platform.

  • the microsoft 2017 conversational Speech Recognition System
    International Conference on Acoustics Speech and Signal Processing, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas Stolcke
    Abstract:

    We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.

  • The Microsoft 2017 Conversational Speech Recognition System
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.

  • the microsoft 2016 conversational Speech Recognition System
    arXiv: Computation and Language, 2016
    Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.

  • using mlp features in sri s conversational Speech Recognition System
    Conference of the International Speech Communication Association, 2005
    Co-Authors: Qifeng Zhu, Andreas Stolcke, Barry Y Chen, Nelson Morgan
    Abstract:

    We describe the development of a Speech Recognition System for conversational telephone Speech (CTS) that incorporates acoustic features estimated by multilayer perceptrons (MLP). The acoustic features are based on frame-level phone posterior probabilities, obtained by merging two different MLP estimators, one based on PLP-Tandem features, the other based on hidden activation TRAPs (HATs) features. This paper focuses on the challenges arising when incorporating these nonstandard features into a full-scale Speech-to-text (STT) System, as used by SRI in the Fall 2004 DARPA STT evaluations. First, we developed a series of time-saving techniques for training feature MLPs on 1800 hours of Speech. Second, we investigated which components of a multipass, multi-front-end Recognition System are most profitably augmented with MLP features for best overall performance. The final System obtained achieved a 2% absolute (10% relative) WER reduction over a comparable baseline System that did not include Tandem/HATs MLP features.

Xuedong Huang - One of the best experts on this subject based on the ideXlab platform.

  • the microsoft 2017 conversational Speech Recognition System
    International Conference on Acoustics Speech and Signal Processing, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas Stolcke
    Abstract:

    We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.

  • The Microsoft 2017 Conversational Speech Recognition System
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018
    Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.

  • the microsoft 2016 conversational Speech Recognition System
    arXiv: Computation and Language, 2016
    Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey Zweig
    Abstract:

    We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.

  • Speech Recognition System and method employing data compression
    Journal of the Acoustical Society of America, 1997
    Co-Authors: Xuedong Huang, Shenzhi Zhang
    Abstract:

    A data compression System greatly compresses the stored data used by a Speech Recognition System employing hidden Markov models (HMM). The Speech Recognition System vector quantizes the acoustic space spoken by humans by dividing it into a predetermined number of acoustic features that are stored as codewords in a vector quantization (output probability) table or codebook. For each spoken word, the Speech Recognition System calculates an output probability value for each codeword, the output probability value representing an estimated probability that the word will be spoken using the acoustic feature associated with the codeword. The probability values are stored in an output probability table indexed by each codeword and by each word in a vocabulary. The output probability table is arranged to allow compression of the probability of values associated with each codeword based on other probability values associated with the same codeword, thereby compressing the stored output probability. By compressing the probability values associated with each codeword separate from the probability values associated with other codewords, the Speech Recognition System can recognize spoken words without having to decompress the entire output probability table. In a preferred embodiment, additional compression is achieved by quantizing the probability values into 16 buckets with an equal number of probability values in each bucket. By quantizing the probability values into buckets, additional redundancy is added to the output probability table, which allows the output probability table to be additionally compressed.

  • an overview of the sphinx ii Speech Recognition System
    Human Language Technology, 1993
    Co-Authors: Xuedong Huang, M.-y. Hwang, Fileno Alleva, Ronald Rosenfeld
    Abstract:

    In the past year at Carnegie Mellon steady progress has been made in the area of acoustic and language modeling. The result has been a dramatic reduction in Speech Recognition errors in the SPHINX-II System. In this paper, we review SPHINX-II and summarize our recent efforts on improved Speech Recognition. Recently SPHINX-II achieved the lowest error rate in the November 1992 DARPA evaluations. For 5000-word, speaker-independent, continuous, Speech Recognition, the error rate was reduced to 5%.