The Experts below are selected from a list of 88806 Experts worldwide ranked by ideXlab platform
Li Deng - One of the best experts on this subject based on the ideXlab platform.
-
discriminative pronounciation learning using phonetic decoder and minimum Classification Error criterion
International Conference on Acoustics Speech and Signal Processing, 2009Co-Authors: Oriol Vinyals, Li Deng, Alex AceroAbstract:In this paper, we report our recent research aimed at improving the pronunciation-modeling component of a speech recognition system designed for mobile voice search. Our new discriminative learning technique overcomes the limitation of the traditional ways of introducing alternative pronunciations that often enlarge confusability across different lexical items. Instead, we make use of a phonetic recognizer to generate pronunciation candidates, which are then evaluated and selected using the global minimum-Classification-Error measure, guaranteeing a reduction of the training-set Error rate after introducing alternative pronunciations. A maximum entropy approach is subsequently used to learn the weight parameters of the selected pronunciation candidates. Our experimental results demonstrate the effectiveness of the discriminative pronunciation learning technique in a real-world speech recognition task where pronunciation of business names presents special difficulty for high-accuracy speech recognition.
-
large margin minimum Classification Error training a theoretical risk minimization perspective
Computer Speech & Language, 2008Co-Authors: Li Deng, Alejandro AceroAbstract:Large-margin discriminative training of hidden Markov models has received significant attention recently. A natural and interesting question is whether the existing discriminative training algorithms can be extended directly to embed the concept of margin. In this paper, we give this question an affirmative answer by showing that the sigmoid bias in the conventional minimum Classification Error (MCE) training can be interpreted as a soft margin. We justify this claim from a theoretical Classification risk minimization perspective where the loss function associated with a non-zero sigmoid bias is shown to include not only empirical Error rates but also a margin-bound risk. Based on this perspective, we propose a practical optimization strategy that adjusts the margin (sigmoid bias) incrementally in the MCE training process so that a desirable balance between the empirical Error rates on the training set and the margin can be achieved. We call this modified MCE training process large-margin minimum Classification Error (LM-MCE) training to differentiate it from the conventional MCE. Speech recognition experiments have been carried out on two tasks. First, in the TIDIGITS recognition task, LM-MCE outperforms the state-of-the-art MCE method with 17% relative digit-Error reduction and 19% relative string-Error reduction. Second, on the Microsoft internal large vocabulary telephony speech recognition task (with 2000h of training data and 120K words in the vocabulary), significant recognition accuracy improvement is achieved, demonstrating that our formulation of LM-MCE can be successfully scaled up and applied to large-scale speech recognition tasks.
-
phone discriminating minimum Classification Error p mce training for phonetic recognition
Conference of the International Speech Communication Association, 2007Co-Authors: Qian Qian, Li DengAbstract:In this paper, we report a study on performance comparisons of discriminative training methods for phone recognition using the TIMIT database. We propose a new method of phonediscriminating minimum Classification Error (P-MCE), which performs MCE training at the sub-string or phone level instead of at the traditional string level. Aiming at minimizing the phone recognition Error rate, P-MCE nevertheless takes advantage of the well-known, efficient training routine derived from the conventional string-based MCE, using specially constructed one-best lists selected from phone lattices. Extensive investigations and comparisons are conducted between the PMCE and other discriminative training methods including maximum mutual information (MMI), minimum phone or word Error (MPE/MWE), and the other two MCE methods. The P-MCE outperforms most of experimented approaches on the standard TIMIT database in terms of the continuous phonetic recognition accuracy. P-MCE achieves comparable results with the MPE method which also aims at reducing phone-level recognition Errors.
-
large margin minimum Classification Error training for large scale speech recognition tasks
International Conference on Acoustics Speech and Signal Processing, 2007Co-Authors: Li Deng, Alejandro AceroAbstract:Recently, we have developed a novel discriminative training method named large-margin minimum Classification Error (LM-MCE) training that incorporates the idea of discriminative margin into the conventional minimum Classification Error (MCE) training method. In our previous work, this novel approach was formulated specifically for the MCE training using the sigmoid loss function and its effectiveness was demonstrated on the TIDIGITS task alone. In this paper two additional contributions are made. First, we formulate LM-MCE as a Bayes risk minimization problem whose loss function not only includes empirical Error rates but also a margin-bound risk. This new formulation allows us to extend the same technique to a wide variety of MCE based training. Second, we have successfully applied LM-MCE training approach to the Microsoft internal large vocabulary telephony speech recognition task (with 2000 hours of training data and 120K of vocabulary) and achieved significant recognition accuracy improvement across-the-board. To our best knowledge, this is the first time that the large-margin approach is demonstrated to be successful in large-scale speech recognition tasks.
-
speech trajectory discrimination using the minimum Classification Error learning
IEEE Transactions on Speech and Audio Processing, 1998Co-Authors: R Chengalvarayan, Li DengAbstract:In this paper, we extend the maximum likelihood (ML) training algorithm to the minimum Classification Error (MCE) training algorithm for discriminatively estimating the state-dependent polynomial coefficients in the stochastic trajectory model or the trended hidden Markov model (HMM) originally proposed in Deng (1992). The main motivation of this extension is the new model space for smoothness-constrained, state-bound speech trajectories associated with the trended HMM, contrasting the conventional, stationary-state HMM, which describes only the piecewise-constant "degraded trajectories" in the observation data. The discriminative training implemented for the trended HMM has the potential to utilize this new, constrained model space, thereby providing stronger power to disambiguate the observational trajectories generated from nonstationary sources corresponding to different speech classes. Phonetic Classification results are reported which demonstrate consistent performance improvements with use of the MCE-trained trended HMM both over the regular ML-trained trended HMM and over the MCE-trained stationary-state HMM.
Wu Chou - One of the best experts on this subject based on the ideXlab platform.
-
Discriminant-function-based minimum recognition Error rate pattern-recognition approach to speech recognition
Proceedings of the IEEE, 2000Co-Authors: Wu ChouAbstract:A discriminant function-based minimum recognition Error rate pattern recognition approach is described and studied for various applications in speech processing. This approach departs from the conventional paradigm, which links a Classification/recognition task to the problem of distribution estimation. Instead, it takes a discriminant function based statistical pattern recognition approach. The suitability of this approach for Classification Error rate minimization is established through a special loss function. It is meaningful even when the model correctness assumption is known to be not valid. We study the theoretical basis of this approach and compare it with various criteria used in speech recognition. We differentiate the method of classifier design by way of distribution estimation and the discriminant function methods of minimizing Classification Error rate, based on the fact that in many realistic applications, such as speech recognition, the true distribution form of the source is rarely known precisely, and without model correctness assumption, the classical optimality theory of the distribution estimation approach cannot be applied directly. We discuss issues in this new classifier design paradigm and present various extensions of this approach to classifier design applications in speech processing.
-
Minimum Classification Error rate methods for speech recognition
IEEE Transactions on Speech and Audio Processing, 1997Co-Authors: Biing-hwang Fred Juang, Wu ChouAbstract:A critical component in the pattern matching approach to speech recognition is the training algorithm, which aims at producing typical (reference) patterns or models for accurate pattern comparison. In this paper, we discuss the issue of speech recognizer training from a broad perspective with root in the classical Bayes decision theory. We differentiate the method of classifier design by way of distribution estimation and the discriminative method of minimizing Classification Error rate based on the fact that in many realistic applications, such as speech recognition, the real signal distribution form is rarely known precisely. We argue that traditional methods relying on distribution estimation are suboptimal when the assumed distribution form is not the true one, and that “optimality” in distribution estimation does not automatically translate into “optimality” in classifier design. We compare the two different methods in the context of hidden Markov modeling for speech recognition. We show the superiority of the minimum Classification Error (MCE) method over the distribution estimation method by providing the results of several key speech recognition experiments. In general, the MCE method provides a significant reduction of recognition Error rate
Alejandro Acero - One of the best experts on this subject based on the ideXlab platform.
-
large margin minimum Classification Error training a theoretical risk minimization perspective
Computer Speech & Language, 2008Co-Authors: Li Deng, Alejandro AceroAbstract:Large-margin discriminative training of hidden Markov models has received significant attention recently. A natural and interesting question is whether the existing discriminative training algorithms can be extended directly to embed the concept of margin. In this paper, we give this question an affirmative answer by showing that the sigmoid bias in the conventional minimum Classification Error (MCE) training can be interpreted as a soft margin. We justify this claim from a theoretical Classification risk minimization perspective where the loss function associated with a non-zero sigmoid bias is shown to include not only empirical Error rates but also a margin-bound risk. Based on this perspective, we propose a practical optimization strategy that adjusts the margin (sigmoid bias) incrementally in the MCE training process so that a desirable balance between the empirical Error rates on the training set and the margin can be achieved. We call this modified MCE training process large-margin minimum Classification Error (LM-MCE) training to differentiate it from the conventional MCE. Speech recognition experiments have been carried out on two tasks. First, in the TIDIGITS recognition task, LM-MCE outperforms the state-of-the-art MCE method with 17% relative digit-Error reduction and 19% relative string-Error reduction. Second, on the Microsoft internal large vocabulary telephony speech recognition task (with 2000h of training data and 120K words in the vocabulary), significant recognition accuracy improvement is achieved, demonstrating that our formulation of LM-MCE can be successfully scaled up and applied to large-scale speech recognition tasks.
-
large margin minimum Classification Error training for large scale speech recognition tasks
International Conference on Acoustics Speech and Signal Processing, 2007Co-Authors: Li Deng, Alejandro AceroAbstract:Recently, we have developed a novel discriminative training method named large-margin minimum Classification Error (LM-MCE) training that incorporates the idea of discriminative margin into the conventional minimum Classification Error (MCE) training method. In our previous work, this novel approach was formulated specifically for the MCE training using the sigmoid loss function and its effectiveness was demonstrated on the TIDIGITS task alone. In this paper two additional contributions are made. First, we formulate LM-MCE as a Bayes risk minimization problem whose loss function not only includes empirical Error rates but also a margin-bound risk. This new formulation allows us to extend the same technique to a wide variety of MCE based training. Second, we have successfully applied LM-MCE training approach to the Microsoft internal large vocabulary telephony speech recognition task (with 2000 hours of training data and 120K of vocabulary) and achieved significant recognition accuracy improvement across-the-board. To our best knowledge, this is the first time that the large-margin approach is demonstrated to be successful in large-scale speech recognition tasks.
Shigeru Katagiri - One of the best experts on this subject based on the ideXlab platform.
-
Robust and Efficient Pattern Classification using Large Geometric Margin Minimum Classification Error Training
Journal of Signal Processing Systems, 2014Co-Authors: Hideyuki Watanabe, Shigeru Katagiri, Tsukasa Ohashi, Miho Ohsaki, Shigeki Matsuda, Hideki KashiokaAbstract:Recently, one of the standard discriminative training methods for pattern classifier design, i.e., Minimum Classification Error (MCE) training, has been revised, and its new version is called Large Geometric Margin Minimum Classification Error (LGM-MCE) training. It is formulated by replacing a conventional misClassification measure, which is equivalent to the so-called functional margin, with a geometric margin that represents the geometric distance between an estimated class boundary and its closest training pattern sample. It seeks the status of the trainable classifier parameters that simultaneously correspond to the minimum of the empirical average Classification Error count loss and the maximum of the geometric margin. Experimental evaluations showed the fundamental utility of LGM-MCE training. However, to increase its effectiveness, this new training required careful setting for hyperparameters, especially the smoothness degree of the smooth Classification Error count loss. Exploring the smoothness degree usually requires many trial-and-Error repetitions of training and testing, and such burdensome repetition does not necessarily lead to an optimal smoothness setting. To alleviate this problem and further increase the effect of geometric margin employment, we apply in this paper a new idea that automatically determines the loss smoothness of LGM-MCE training. We first introduce a new formalization of it using the Parzen estimation of Error count risk and formalize LGM-MCE training that incorporates a mechanism of automatic loss smoothness determination. Importantly, the geometric-margin-based misClassification measure adopted in LGM-MCE training is directly linked with the geometric margin in a pattern sample space. Based on this relation, we also prove that loss smoothness affects the production of virtual samples along the estimated class boundaries in pattern sample space. Finally, through experimental evaluations and in comparisons with other training methods, we elaborate the characteristics of LGM-MCE training and its new function that automatically determines an appropriate loss smoothness degree.
-
Discriminative Training for Large-Vocabulary Speech Recognition Using Minimum Classification Error
IEEE Transactions on Audio Speech and Language Processing, 2007Co-Authors: Erik Mcdermott, Timothy J. Hazen, Jonathan Roux, Atsushi Nakamura, Shigeru KatagiriAbstract:The minimum Classification Error (MCE) framework for discriminative training is a simple and general formalism for directly optimizing recognition accuracy in pattern recognition problems. The framework applies directly to the optimization of hidden Markov models (HMMs) used for speech recognition problems. However, few if any studies have reported results for the application of MCE training to large-vocabulary, contin- uous-speech recognition tasks. This article reports significant gains in recognition performance and model compactness as a result of discriminative training based onMCEtraining applied to HMMs, in the context of three challenging large-vocabulary (up to 100 k word) speech recognition tasks: the Corpus of Spontaneous Japanese lecture speech transcription task, a telephone-based name recognition task, and the MIT JUPITER telephone-based conversational weather information task. On these tasks, starting from maximum likelihood (ML) baselines, MCE training yielded relative reductions in word Error ranging from 7% to 20%. Furthermore, this paper evaluates the use of different methods for optimizing the MCE criterion function, as well as the use of pre- computed recognition lattices to speed up training. An overview of the MCE framework is given, with an emphasis on practical implementation issues.
-
a derivation of minimum Classification Error from the theoretical Classification risk using parzen estimation
Computer Speech & Language, 2004Co-Authors: Erik Mcdermott, Shigeru KatagiriAbstract:The minimum Classification Error (MCE) framework is an approach to discriminative training for pattern recognition that explicitly incorporates a smoothed version of Classification performance into the recognizer design criterion. Many studies have confirmed the effectiveness of MCE for speech recognition. In this article, we present a theoretical analysis of the smoothness of the MCE loss function. Specifically, we show that the MCE criterion function is equivalent to a Parzen window-based estimate of the theoretical Classification risk. In this analysis, each training token is mapped to the center of a Parzen kernel in the domain of a suitably defined random variable. The kernels are summed to produce a density estimate; this estimate in turn can easily be integrated over the domain of incorrect Classifications, yielding the risk estimate. The expression of risk for each kernel corresponds directly to the usual MCE loss function. The specific form of the Parzen window corresponds to the specific form of the MCE loss function. The derivation presented here shows that the smooth MCE loss function, far from being an ad-hoc approximation of the true Error, can be seen as the direct consequence of using a well-understood type of smoothing, Parzen estimation, to estimate the theoretical risk from a finite training set. This analysis provides a novel link between the MCE empirical cost measured on a finite training set and the theoretical Classification risk.
Euisun Choi - One of the best experts on this subject based on the ideXlab platform.
-
estimation of Classification Error based on the bhattacharyya distance for multimodal data
International Geoscience and Remote Sensing Symposium, 2001Co-Authors: Euisun Choi, Chulhee LeeAbstract:In this paper, we investigate the possibility of Error estimation based on the Bhattacharyya distance for multimodal data. Assuming multimodal data can be approximated as a mixture of several classes that has the Gaussian distribution, we try to find the empirical relationship between the Bhattacharyya distance and the Classification Error for multimodal data. Experimental results with remotely sensed data showed that there exists a strong relationship and that it is possible to predict the Classification Error using the Bhattacharyya distance for multimodal data.
-
feature extraction based on the bhattacharyya distance for multimodal data
International Geoscience and Remote Sensing Symposium, 2001Co-Authors: Euisun Choi, Chulhee LeeAbstract:In this paper, we propose a feature extraction method based on the Bhattacharyya distance for multimodal data. First, we estimate the Classification Error based on the Bhattacharyya distance between two multimodal classes that are approximated by a finite mixture of Gaussian distributions. Then we extract the features that minimize the estimated Classification Error. In order to find such features, we explore two search methods: sequential search and global search. Experiments show that the proposed feature extraction algorithm shows promising results.
-
Feature extraction based on the Bhattacharyya distance
IGARSS 2000. IEEE 2000 International Geoscience and Remote Sensing Symposium. Taking the Pulse of the Planet: The Role of Remote Sensing in Managing t, 2000Co-Authors: Euisun ChoiAbstract:The authors propose a feature extraction method based on the Bhattacharyya distance. Recently, it has been reported that an accurate estimation of Classification Error is possible using the Bhattacharyya distance. In the proposed method, the authors try to find feature vectors that minimize the estimated Classification Error of Gaussian ML classifier. In order to find such feature vectors, they start with arbitrary initial feature vectors and update them using two optimization techniques: sequential search and global search. Since they use the Error estimation equation for updating feature vectors, the search time can be reduced significantly. They first apply the algorithm to two class problems and extend it to multiclass problems. Experimental results show that the proposed feature extraction algorithm compares favorably with conventional feature extraction algorithms.
-
bayes Error evaluation of the gaussian ml classifier
IEEE Transactions on Geoscience and Remote Sensing, 2000Co-Authors: Euisun ChoiAbstract:The authors investigate the relationship between the Classification Error and the Bhattacharyya distance of two normally distributed classes and propose a new equation that provides an accurate Error estimation for the Gaussian ML classifier. With the Error estimation equation, it is possible to estimate the Classification Error within a 1-2% margin.
-
feature extraction method using the bhattacharyya distance
Journal of the Institute of Electronics Engineers of Korea, 2000Co-Authors: Euisun Choi, Chulhee LeeAbstract:In pattern Classification, the Bhattacharyya distance has been used as a class separability measure. Furthemore, it is recently reported that the Bhattacharyya distance can be used to estimate Error of Gaussian ML classifier within 1-2% margin. In this paper, we propose a feature extraction method utilizing the Bhattacharyya distance. In the proposed method, we first predict the Classification Error with the Error estimation equation based on the Bhauacharyya distance. Then we find the feature vector that minimizes the Classification Error using two search algorithms: sequential search and global search. Experimental reslts show that the proposed method compares favorably with conventional feature extraction methods. In addition, it is possible to determine how man, feature vectors arc needed for achieving the same Classification accuracy as in the original space.