The Experts below are selected from a list of 49827 Experts worldwide ranked by ideXlab platform
Fileno Alleva - One of the best experts on this subject based on the ideXlab platform.
-
the microsoft 2017 conversational Speech Recognition System
International Conference on Acoustics Speech and Signal Processing, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas StolckeAbstract:We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.
-
The Microsoft 2017 Conversational Speech Recognition System
ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey ZweigAbstract:We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.
-
an overview of the sphinx ii Speech Recognition System
Human Language Technology, 1993Co-Authors: Xuedong Huang, M.-y. Hwang, Fileno Alleva, Ronald RosenfeldAbstract:In the past year at Carnegie Mellon steady progress has been made in the area of acoustic and language modeling. The result has been a dramatic reduction in Speech Recognition errors in the SPHINX-II System. In this paper, we review SPHINX-II and summarize our recent efforts on improved Speech Recognition. Recently SPHINX-II achieved the lowest error rate in the November 1992 DARPA evaluations. For 5000-word, speaker-independent, continuous, Speech Recognition, the error rate was reduced to 5%.
-
the sphinx ii Speech Recognition System an overview
Computer Speech & Language, 1992Co-Authors: Xuedong Huang, Fileno Alleva, Mei Hwang, Ronald RosenfeldAbstract:In order for Speech recognizers to deal with increased task perplexity, speaker variation, and environment variation, improved Speech Recognition is critical. Steady progress has been made along these three dimensions at Carnegie Mellon. In this paper, we review the SPHINX-II Speech Recognition System and summarize our recent efforts on improved Speech Recognition.
James Droppo - One of the best experts on this subject based on the ideXlab platform.
-
the microsoft 2017 conversational Speech Recognition System
International Conference on Acoustics Speech and Signal Processing, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas StolckeAbstract:We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.
-
The Microsoft 2017 Conversational Speech Recognition System
ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey ZweigAbstract:We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.
-
the microsoft 2016 conversational Speech Recognition System
arXiv: Computation and Language, 2016Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey ZweigAbstract:We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
-
efficient on line acoustic environment estimation for fcdcn in a continuous Speech Recognition System
International Conference on Acoustics Speech and Signal Processing, 2001Co-Authors: James Droppo, Alejandro Acero, Li DengAbstract:There exists a number of cepstral de-noising algorithms which perform quite well when trained and tested under similar acoustic environments, but degrade quickly under mismatched conditions. We present two key results that make these algorithms practical in real noise environments, with the ability to adapt to different acoustic environments over time. First, we show that it is possible to leverage the existing de-noising computations to estimate the acoustic environment on-line and in real time. Second, we show that it is not necessary to collect large amounts of training data in each environment-clean data with artificial mixing is sufficient. When this new method is used as a pre-processing stage to a large vocabulary Speech Recognition System, it can be made robust to a wide variety of acoustic environments. With synthetic training data, we are able to reduce the word error rate by 27%.
George Saon - One of the best experts on this subject based on the ideXlab platform.
-
The IBM 2016 English conversational telephone Speech Recognition System
Proceedings of the Annual Conference of the International Speech Communication Association INTERSPEECH, 2016Co-Authors: George Saon, Tom Sercu, Steven Rennie, Hong Kwang J. KuoAbstract:We describe a collection of acoustic and language modeling techniques that lowered the word error rate of our English conversational telephone LVCSR System to a record 6.6% on the Switchboard subset of the Hub5 2000 evaluation testset. On the acoustic side, we use a score fusion of three strong models: recurrent nets with maxout activations, very deep convolutional nets with 3x3 kernels, and bidirectional long short-term memory nets which operate on FMLLR and i-vector features. On the language modeling side, we use an updated model "M" and hierarchical neural network LMs.
-
the ibm 2015 english conversational telephone Speech Recognition System
arXiv: Computation and Language, 2015Co-Authors: George Saon, Steven J Rennie, Michael PichenyAbstract:We describe the latest improvements to the IBM English conversational telephone Speech Recognition System. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result.
-
the ibm 2015 english conversational telephone Speech Recognition System
Conference of the International Speech Communication Association, 2015Co-Authors: George Saon, Steven J Rennie, Michael PichenyAbstract:We describe the latest improvements to the IBM English conversational telephone Speech Recognition System. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. Index Terms: recurrent neural networks, convolutional neural networks, conversational Speech Recognition
Andreas Stolcke - One of the best experts on this subject based on the ideXlab platform.
-
the microsoft 2017 conversational Speech Recognition System
International Conference on Acoustics Speech and Signal Processing, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas StolckeAbstract:We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.
-
The Microsoft 2017 Conversational Speech Recognition System
ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey ZweigAbstract:We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.
-
the microsoft 2016 conversational Speech Recognition System
arXiv: Computation and Language, 2016Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey ZweigAbstract:We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
-
using mlp features in sri s conversational Speech Recognition System
Conference of the International Speech Communication Association, 2005Co-Authors: Qifeng Zhu, Andreas Stolcke, Barry Y Chen, Nelson MorganAbstract:We describe the development of a Speech Recognition System for conversational telephone Speech (CTS) that incorporates acoustic features estimated by multilayer perceptrons (MLP). The acoustic features are based on frame-level phone posterior probabilities, obtained by merging two different MLP estimators, one based on PLP-Tandem features, the other based on hidden activation TRAPs (HATs) features. This paper focuses on the challenges arising when incorporating these nonstandard features into a full-scale Speech-to-text (STT) System, as used by SRI in the Fall 2004 DARPA STT evaluations. First, we developed a series of time-saving techniques for training feature MLPs on 1800 hours of Speech. Second, we investigated which components of a multipass, multi-front-end Recognition System are most profitably augmented with MLP features for best overall performance. The final System obtained achieved a 2% absolute (10% relative) WER reduction over a comparable baseline System that did not include Tandem/HATs MLP features.
Xuedong Huang - One of the best experts on this subject based on the ideXlab platform.
-
the microsoft 2017 conversational Speech Recognition System
International Conference on Acoustics Speech and Signal Processing, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Xuedong Huang, Andreas StolckeAbstract:We describe the latest version of Microsoft's conversational Speech Recognition System for the Switchboard and CallHome domains. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby acoustic model posteriors are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added another language model rescoring step following the confusion network combination. The resulting System yields a 5.1% word error rate on the NIST 2000 Switchboard test set, and 9.8% on the CallHome subset.
-
The Microsoft 2017 Conversational Speech Recognition System
ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2018Co-Authors: Wayne Xiong, Fileno Alleva, James Droppo, Michael L. Seltzer, Andreas Stolcke, Lingfeng Wu, Frank Seide, Dong Yu, Xuedong Huang, Geoffrey ZweigAbstract:We describe the 2017 version of Microsoft's conversational Speech Recognition System, in which we update our 2016 System with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard Speech Recognition task. The System adds a CNN-BLSTM acoustic model to the set of model architectures we combined previously, and includes character-based and dialog session aware LSTM language models in rescoring. For System combination we adopt a two-stage approach, whereby subsets of acoustic models are first combined at the senone/frame level, followed by a word-level voting via confusion networks. We also added a confusion network rescoring step after System combination. The resulting System yields a 5.1\% word error rate on the 2000 Switchboard evaluation set.
-
the microsoft 2016 conversational Speech Recognition System
arXiv: Computation and Language, 2016Co-Authors: Wayne Xiong, James Droppo, Michael L. Seltzer, Andreas Stolcke, Frank Seide, Xuedong Huang, Geoffrey ZweigAbstract:We describe Microsoft's conversational Speech Recognition System, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard Recognition task. Inspired by machine learning ensemble techniques, the System uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based System combination provide a 20% boost. The best single System uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined System has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
-
Speech Recognition System and method employing data compression
Journal of the Acoustical Society of America, 1997Co-Authors: Xuedong Huang, Shenzhi ZhangAbstract:A data compression System greatly compresses the stored data used by a Speech Recognition System employing hidden Markov models (HMM). The Speech Recognition System vector quantizes the acoustic space spoken by humans by dividing it into a predetermined number of acoustic features that are stored as codewords in a vector quantization (output probability) table or codebook. For each spoken word, the Speech Recognition System calculates an output probability value for each codeword, the output probability value representing an estimated probability that the word will be spoken using the acoustic feature associated with the codeword. The probability values are stored in an output probability table indexed by each codeword and by each word in a vocabulary. The output probability table is arranged to allow compression of the probability of values associated with each codeword based on other probability values associated with the same codeword, thereby compressing the stored output probability. By compressing the probability values associated with each codeword separate from the probability values associated with other codewords, the Speech Recognition System can recognize spoken words without having to decompress the entire output probability table. In a preferred embodiment, additional compression is achieved by quantizing the probability values into 16 buckets with an equal number of probability values in each bucket. By quantizing the probability values into buckets, additional redundancy is added to the output probability table, which allows the output probability table to be additionally compressed.
-
an overview of the sphinx ii Speech Recognition System
Human Language Technology, 1993Co-Authors: Xuedong Huang, M.-y. Hwang, Fileno Alleva, Ronald RosenfeldAbstract:In the past year at Carnegie Mellon steady progress has been made in the area of acoustic and language modeling. The result has been a dramatic reduction in Speech Recognition errors in the SPHINX-II System. In this paper, we review SPHINX-II and summarize our recent efforts on improved Speech Recognition. Recently SPHINX-II achieved the lowest error rate in the November 1992 DARPA evaluations. For 5000-word, speaker-independent, continuous, Speech Recognition, the error rate was reduced to 5%.