The Experts below are selected from a list of 53259 Experts worldwide ranked by ideXlab platform
Maria E Smith - One of the best experts on this subject based on the ideXlab platform.
-
recent improvements to the ibm trainable speech synthesis system
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Ellen Marie Eide, Andy Aaron, Raimo Bakis, R Cohen, Robert E Donovan, Wael Hamza, T Mathes, Michael Picheny, M Polkosky, Maria E SmithAbstract:In this paper we describe the current status of the trainable text-to-speech system at IBM. Recent algorithmic and database changes to the system have led to significant gains in the output quality. On the algorithms side, we have introduced statistical models for predicting pitch and duration targets which replace the rule-based target generation previously employed. Additionally, we have changed the cost function and the search strategy, introduced a post-search pitch smoothing algorithm, and improved our method of preselection. Through the combined data and algorithmic contributions, we have been able to significantly improve (p < 0.0001) the Mean Opinion Score (MOS) of our female voice, from 3.68 to 4.85 when heard over loudspeakers and to 5.42 when heard over the telephone (seven point scale).
J L Crebouw - One of the best experts on this subject based on the ideXlab platform.
-
objective evaluation of hmm based speech synthesis system using kullback leibler divergence
Conference of the International Speech Communication Association, 2014Co-Authors: Congthanh Do, Marc Evrard, A Leman, Albert Rilliard, Christophe Dalessandro, J L CrebouwAbstract:In this paper, we propose a new objective evaluation method for hidden Markov model (HMM)-based speech synthesis using Kullback-Leibler divergence (KLD). The KLD is used to measure the difference between the probability density functions (PDFs) of the acoustic feature vectors extracted from natural training and synthetic speech data. For the evaluation, Gaussian mixture model (GMM) is used to model the distribution of acoustic feature vectors, including the fundamental frequency (F0). Continuous F0, obtained with linear interpolation, is used in the evaluation. In essence, the KLD is the expectation of the logarithmic difference between the likelihoods calculated on training and synthetic speech. This likelihood difference is appropriate to characterize the quality of a HMMbased speech synthesis system in generating synthetic speech using a maximum likelihood criterion. The objective evaluation is tested with 3 different HMM-based speech synthesis systems which use multi-space distribution (MSD) to model discontinuous F0. These systems are trained on a common speech corpus in French. We propose an index to evaluate HMM-based speech synthesis system which takes into account the relative variation of the KLDs on test sets of synthetic and natural speech. This index correlates inversely with the result of the MOS (Mean Opinion Score) perceptual test.
Ellen Marie Eide - One of the best experts on this subject based on the ideXlab platform.
-
recent improvements to the ibm trainable speech synthesis system
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Ellen Marie Eide, Andy Aaron, Raimo Bakis, R Cohen, Robert E Donovan, Wael Hamza, T Mathes, Michael Picheny, M Polkosky, Maria E SmithAbstract:In this paper we describe the current status of the trainable text-to-speech system at IBM. Recent algorithmic and database changes to the system have led to significant gains in the output quality. On the algorithms side, we have introduced statistical models for predicting pitch and duration targets which replace the rule-based target generation previously employed. Additionally, we have changed the cost function and the search strategy, introduced a post-search pitch smoothing algorithm, and improved our method of preselection. Through the combined data and algorithmic contributions, we have been able to significantly improve (p < 0.0001) the Mean Opinion Score (MOS) of our female voice, from 3.68 to 4.85 when heard over loudspeakers and to 5.42 when heard over the telephone (seven point scale).
Congthanh Do - One of the best experts on this subject based on the ideXlab platform.
-
objective evaluation of hmm based speech synthesis system using kullback leibler divergence
Conference of the International Speech Communication Association, 2014Co-Authors: Congthanh Do, Marc Evrard, A Leman, Albert Rilliard, Christophe Dalessandro, J L CrebouwAbstract:In this paper, we propose a new objective evaluation method for hidden Markov model (HMM)-based speech synthesis using Kullback-Leibler divergence (KLD). The KLD is used to measure the difference between the probability density functions (PDFs) of the acoustic feature vectors extracted from natural training and synthetic speech data. For the evaluation, Gaussian mixture model (GMM) is used to model the distribution of acoustic feature vectors, including the fundamental frequency (F0). Continuous F0, obtained with linear interpolation, is used in the evaluation. In essence, the KLD is the expectation of the logarithmic difference between the likelihoods calculated on training and synthetic speech. This likelihood difference is appropriate to characterize the quality of a HMMbased speech synthesis system in generating synthetic speech using a maximum likelihood criterion. The objective evaluation is tested with 3 different HMM-based speech synthesis systems which use multi-space distribution (MSD) to model discontinuous F0. These systems are trained on a common speech corpus in French. We propose an index to evaluate HMM-based speech synthesis system which takes into account the relative variation of the KLDs on test sets of synthetic and natural speech. This index correlates inversely with the result of the MOS (Mean Opinion Score) perceptual test.
Wael Hamza - One of the best experts on this subject based on the ideXlab platform.
-
recent improvements to the ibm trainable speech synthesis system
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Ellen Marie Eide, Andy Aaron, Raimo Bakis, R Cohen, Robert E Donovan, Wael Hamza, T Mathes, Michael Picheny, M Polkosky, Maria E SmithAbstract:In this paper we describe the current status of the trainable text-to-speech system at IBM. Recent algorithmic and database changes to the system have led to significant gains in the output quality. On the algorithms side, we have introduced statistical models for predicting pitch and duration targets which replace the rule-based target generation previously employed. Additionally, we have changed the cost function and the search strategy, introduced a post-search pitch smoothing algorithm, and improved our method of preselection. Through the combined data and algorithmic contributions, we have been able to significantly improve (p < 0.0001) the Mean Opinion Score (MOS) of our female voice, from 3.68 to 4.85 when heard over loudspeakers and to 5.42 when heard over the telephone (seven point scale).