The Experts below are selected from a list of 12903 Experts worldwide ranked by ideXlab platform
Mukund Padmanabhan - One of the best experts on this subject based on the ideXlab platform.
-
Minimum Bayes error feature selection for Continuous Speech Recognition
2020Co-Authors: George Saon, Mukund PadmanabhanAbstract:Abstract We consider the problem of designing a linear transformation () E lR Px n, of rank p ~ n, which projects the features of a classifier x E lR n onto y = ()x E lR P such as to achieve minimum Bayes error (or probability of misclassification). Two avenues will be explored: the first is to maximize the ()-average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of (). While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony Speech Recognition task
-
data driven approach to designing compound words for Continuous Speech Recognition
IEEE Transactions on Speech and Audio Processing, 2001Co-Authors: George Saon, Mukund PadmanabhanAbstract:We present a new approach to deriving compound words from a training corpus. The motivation for making compound words is because under some assumptions, Speech Recognition errors occur less frequently in longer words. Furthermore, they also enable more accurate modeling of pronunciation variability at the boundary between adjacent words in a Continuously spoken utterance. We introduce a measure based on the product between the direct and the reverse bigram probability of a pair of words for finding candidate pairs in order to create compound words. Our experimental results show that by augmenting both the acoustic vocabulary and the language model with these new tokens, the word Recognition accuracy can be improved by absolute 2.8% (7% relative) on a voice mail Continuous Speech Recognition task. We also compare the proposed measure for selecting compound words with other measures that have been described in the literature.
-
minimum bayes error feature selection for Continuous Speech Recognition
Neural Information Processing Systems, 2000Co-Authors: George Saon, Mukund PadmanabhanAbstract:We consider the problem of designing a linear transformation θ ∈ Rp × n, of rank p ≤ n, which projects the features of a classifier x ∈ Rn onto y = θx ∈ Rp such as to achieve minimum Bayes error (or probability of misclassification). Two avenues will be explored: the first is to maximize the θ-average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of θ. While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony Speech Recognition task.
-
performance of the ibm large vocabulary Continuous Speech Recognition system on the arpa wall street journal task
International Conference on Acoustics Speech and Signal Processing, 1995Co-Authors: Lalit R Bahl, S Balakrishnanaiyer, J R Bellgarda, Martin Franz, Ponani S Gopalakrishnan, David Nahamoo, Miroslav Novak, Mukund Padmanabhan, Michael Picheny, Salim RoukosAbstract:In this paper we discuss various experimental results using our Continuous Speech Recognition system on the Wall Street Journal task. Experiments with different feature extraction methods, varying amounts and type of training data, and different vocabulary sizes are reported.
George Saon - One of the best experts on this subject based on the ideXlab platform.
-
Recent Developments in Large Vocabulary Continuous Speech Recognition
2020Co-Authors: George Saon, Jen-tzung ChienAbstract:Abstract-This paper overviews a series of recent approaches to front-end processing, acoustic modeling, language modeling, and back-end search and system combination which have made contributions for large vocabulary Continuous Speech Recognition (LVCSR) systems. These approaches include the feature transformations, speaker-adaptive features, and discriminative features in front-end processing, the feature-space and modelspace discriminative training, deep neural networks, and speaker adaptation in acoustic modeling, the backoff smoothing, largespan modeling, and model regularization in language modeling, and the system combination, cross-adaptation, and boosting in search and system combination. Some future directions for LVCSR research are also addressed
-
Minimum Bayes error feature selection for Continuous Speech Recognition
2020Co-Authors: George Saon, Mukund PadmanabhanAbstract:Abstract We consider the problem of designing a linear transformation () E lR Px n, of rank p ~ n, which projects the features of a classifier x E lR n onto y = ()x E lR P such as to achieve minimum Bayes error (or probability of misclassification). Two avenues will be explored: the first is to maximize the ()-average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of (). While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony Speech Recognition task
-
Large-vocabulary Continuous Speech Recognition systems: A look at some recent advances
IEEE Signal Processing Magazine, 2012Co-Authors: George Saon, Jen-tzung ChienAbstract:Over the past decade or so, several advances have been made to the design of modern large vocabulary Continuous Speech Recognition (LVCSR) systems to the point where their application has broadened from early speaker dependent dictation systems to speaker-independent automatic broadcast news transcription and indexing, lectures and meetings transcription, conversational telephone Speech transcription, open-domain voice search, medical and legal Speech Recognition, and call center applications, to name a few. The commercial success of these systems is an impressive testimony to how far research in LVCSR has come, and the aim of this article is to describe some of the technological underpinnings of modern systems. It must be said, however, that, despite the commercial success and widespread adoption, the problem of large-vocabulary Speech Recognition is far from being solved: background noise, channel distortions, foreign accents, casual and disfluent Speech, or unexpected topic change can cause automated systems to make egregious Recognition errors. This is because current LVCSR systems are not robust to mismatched training and test conditions and cannot handle context as well as human listeners despite being trained on thousands of hours of Speech and billions of words of text.
-
data driven approach to designing compound words for Continuous Speech Recognition
IEEE Transactions on Speech and Audio Processing, 2001Co-Authors: George Saon, Mukund PadmanabhanAbstract:We present a new approach to deriving compound words from a training corpus. The motivation for making compound words is because under some assumptions, Speech Recognition errors occur less frequently in longer words. Furthermore, they also enable more accurate modeling of pronunciation variability at the boundary between adjacent words in a Continuously spoken utterance. We introduce a measure based on the product between the direct and the reverse bigram probability of a pair of words for finding candidate pairs in order to create compound words. Our experimental results show that by augmenting both the acoustic vocabulary and the language model with these new tokens, the word Recognition accuracy can be improved by absolute 2.8% (7% relative) on a voice mail Continuous Speech Recognition task. We also compare the proposed measure for selecting compound words with other measures that have been described in the literature.
-
minimum bayes error feature selection for Continuous Speech Recognition
Neural Information Processing Systems, 2000Co-Authors: George Saon, Mukund PadmanabhanAbstract:We consider the problem of designing a linear transformation θ ∈ Rp × n, of rank p ≤ n, which projects the features of a classifier x ∈ Rn onto y = θx ∈ Rp such as to achieve minimum Bayes error (or probability of misclassification). Two avenues will be explored: the first is to maximize the θ-average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of θ. While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony Speech Recognition task.
Dennis Norris - One of the best experts on this subject based on the ideXlab platform.
-
Shortlist B: a Bayesian model of Continuous Speech Recognition.
Psychological review, 2008Co-Authors: Dennis Norris, James M McqueenAbstract:A Bayesian model of Continuous Speech Recognition is presented. It is based on Shortlist (D. Norris, 1994; D. Norris, J. M. McQueen, A. Cutler, & S. Butterfield, 1997) and shares many of its key assumptions: parallel competitive evaluation of multiple lexical hypotheses, phonologically abstract prelexical and lexical representations, a feedforward architecture with no online feedback, and a lexical segmentation algorithm based on the viability of chunks of the input as possible words. Shortlist B is radically different from its predecessor in two respects. First, whereas Shortlist was a connectionist model based on interactive-activation principles, Shortlist B is based on Bayesian principles. Second, the input to Shortlist B is no longer a sequence of discrete phonemes; it is a sequence of multiple phoneme probabilities over 3 time slices per segment, derived from the performance of listeners in a large-scale gating study. Simulations are presented showing that the model can account for key findings: data on the segmentation of Continuous Speech, word frequency effects, the effects of mispronunciations on word Recognition, and evidence on lexical involvement in phonemic decision making. The success of Shortlist B suggests that listeners make optimal Bayesian decisions during spoken-word Recognition.
-
Shortlist: a connectionist model of Continuous Speech Recognition
Cognition, 1994Co-Authors: Dennis NorrisAbstract:Previous work has shown a back-propagation network with recurrent connections can successfully model many aspects of human spoken word Recognition (Norris, 1988, 1990, 1992, 1993). However, such networks are unable to revise their decisions in the light of subsequent context. TRACE (McClelland & Elman, 1986), on the other hand, manages to deal appropriately with following context, but only by using a highly implausible architecture that fails to account for some important experimental results. A new model is presented which displays the more desirable properties of each of these models. In contrast to TRACE the new model is entirely bottom-up and can readily perform simulations with vocabularies of tens of thousands of words. © 1994.
Hermann Ney - One of the best experts on this subject based on the ideXlab platform.
-
unsupervised training of acoustic models for large vocabulary Continuous Speech Recognition
IEEE Transactions on Speech and Audio Processing, 2005Co-Authors: Frank Wessel, Hermann NeyAbstract:For large vocabulary Continuous Speech Recognition systems, the amount of acoustic training data is of crucial importance. In the past, large amounts of Speech were thus recorded from various sources and had to be transcribed manually. It is thus desirable to train a recognizer with as little manually transcribed acoustic data as possible. Since untranscribed Speech is available in various forms nowadays, the unsupervised training of a Speech recognizer on recognized transcriptions is studied in this paper. A low-cost recognizer trained with between one and six h of manually transcribed Speech is used to recognize 72 h of untranscribed acoustic data. These transcriptions are then used in combination with a confidence measure to train an improved recognizer. The effect of the confidence measure which is used to detect possible Recognition errors is studied systematically. Finally, the unsupervised training is applied iteratively. Starting with only one h of transcribed acoustic data, a Recognition system is trained fully automatically. With this iterative training procedure, the word error rates are reduced from 71.3% to 38.3% on the Broadcast News'96 evaluation test set and from 65.6% to 29.3% on the Broadcast News'98 evaluation test set. In comparison with an optimized system trained with the manually generated transcriptions of the complete 72 h training corpus, the word error rates increase by 14.3% relative and 18.6% relative, respectively.
-
dynamic programming search for Continuous Speech Recognition
IEEE Signal Processing Magazine, 1999Co-Authors: Hermann Ney, S OrtmannsAbstract:The authors gives a unifying view of the dynamic programming approach to the search problem. They review the search problem from the statistical point-of-view and show how the search space results from the acoustic and language models required by the statistical approach. Starting from the baseline one-pass algorithm using a linear organization of the pronunciation lexicon, they have extended the baseline algorithm toward various dimensions. To handle a large vocabulary, they have shown how the search space can be structured in combination with a lexical prefix tree organization of the pronunciation lexicon. In addition, they have shown how this structure of the search space can be combined with a time-synchronous beam search concept and how the search space can be constructed dynamically during the Recognition process. In particular, to increase the efficiency of the beam search concept, they have integrated the language model look-ahead into the pruning operation. To produce sentence alternatives rather than only the single best sentence, they have extended the search strategy to generate a word graph. Finally, they have reported experimental results on a 64 k-word task that demonstrate the efficiency of the various search concepts presented.
-
the rwth large vocabulary Continuous Speech Recognition system
International Conference on Acoustics Speech and Signal Processing, 1998Co-Authors: Hermann Ney, S Ortmanns, Lutz Welling, Klaus Beulen, Frank WesselAbstract:We present an overview of the RWTH Aachen large vocabulary Continuous Speech recognizer. The recognizer is based on Continuous density hidden Markov models and a time-synchronous left-to-right beam search strategy. Experimental results on the ARPA Wall Street Journal (WSJ) corpus verify the effects of several system components, namely linear discriminant analysis, vocal tract normalization, pronunciation lexicon and cross-word triphones, on the Recognition performance.
-
a word graph algorithm for large vocabulary Continuous Speech Recognition
Computer Speech & Language, 1997Co-Authors: S Ortmanns, Hermann Ney, Xavier L AubertAbstract:Abstract This paper describes a method for the construction of a word graph (or lattice) for large vocabulary, Continuous Speech Recognition. The advantage of a word graph is that a fairly good degree of decoupling between acoustic Recognition at the 10-ms level and the final search at the word level using a complicated language model can be achieved. The word graph algorithm is obtained as an extension of the one-pass beam search strategy using word dependent copies of the word models or lexical trees. The method has been tested successfully on the 20 000-word NAB'94 task (American English, Continuous Speech, 20 000 words, speaker independent) and compared with the integrated method. The experiments show that the word graph density can be reduced to an average number of about 10 word hypotheses, i.e. word edges in the graph, per spoken word with virtually no loss in Recognition performance.
-
Word graphs: an efficient interface between Continuous-Speech Recognition and language understanding
Acoustics, Speech, and Signal Processing, 1993. ICASSP-93., 1993 IEEE International Conference on, 1993Co-Authors: M. Oerder, Hermann NeyAbstract:Word graphs are directed acyclic graphs where each edge is labeled with a word and a score, and each node is labeled with a point in time. Word graphs form an efficient feedforward interface between Continuous-Speech Recognition and linguistic processors. Word graphs with high coverage and modest graph densities can be generated with a computational load comparable with bigram best-sentence Recognition. Results on word graph error rates and word graph densities are presented for the ASL (Architecture Speech/Language) benchmark test.
Piet Wambacq - One of the best experts on this subject based on the ideXlab platform.
-
Template-based Continuous Speech Recognition
IEEE Transactions on Audio, Speech and Language Processing, 2007Co-Authors: Mathias De Wachter, Kris Demuynck, Ronald Cools, Piet Wambacq, Mike Matton, Dirk Van CompernolleAbstract:Despite their known weaknesses, hidden Markov models (HMMs) have been the dominant technique for acoustic modeling in Speech Recognition for over two decades. Still, the advances in the HMM framework have not solved its key problems: it discards information about time dependencies and is prone to overgeneralization. In this paper, we attempt to overcome these problems by relying on straightforward template matching. The basis for the recognizer is the well-known DTW algorithm. However, classical DTW Continuous Speech Recognition results in an explosion of the search space. The traditional top-down search is therefore complemented with a data-driven selection of candidates for DTW alignment. We also extend the DTW framework with a flexible subword unit mechanism and a class sensitive distance measure-two components suggested by state-of-the-art HMM systems. The added flexibility of the unit selection in the template-based framework leads to new approaches to speaker and environment adaptation. The template matching system reaches a performance somewhat worse than the best published HMM results for the Resource Management benchmark, but thanks to complementarity of errors between the HMM and DTW systems, the combination of both leads to a decrease in word error rate with 17% compared to the HMM results
-
data driven example based Continuous Speech Recognition
Conference of the International Speech Communication Association, 2003Co-Authors: Mathias De Wachter, Kris Demuynck, Dirk Van Compernolle, Piet WambacqAbstract:The dominant acoustic modeling methodology based on Hidden Markov Models is known to have certain weaknesses. Partial solutions to these flaws have been presented, but the fundamental problem remains: compression of the data to a compact HMM discards useful information such as time dependencies and speaker information. In this paper, we look at pure example based Recognition as a solution to this problem. By replacing the HMM with the underlying examples, all information in the training data is retained. We show how information about speaker and environment can be used, introducing a new interpretation of adaptation. The basis for the recognizer is the wellknown DTW algorithm, which has often been used for small tasks. However, large vocabulary Speech Recognition introduces new demands, resulting in an explosion of the search space. We show how this problem can be tackled using a data driven approach which selects appropriate Speech examples as candidates for DTW-alignment.