The Experts below are selected from a list of 8685 Experts worldwide ranked by ideXlab platform
Li Deng - One of the best experts on this subject based on the ideXlab platform.
-
adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
IEEE Transactions on Audio Speech and Language Processing, 2007Co-Authors: Li Deng, Hagai Attias, Leo J Lee, Alejandro AceroAbstract:A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent Speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonlinear function is adaptively linearized, and a residual or bias term, which is adaptively trained, is added to the nonlinear function to represent the iteratively reduced piecewise linear approximation error. Details of the piecewise linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard Speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics
-
learning statistically characterized resonance targets in a hidden trajectory model of Speech coarticulation and reduction
Conference of the International Speech Communication Association, 2005Co-Authors: Li Deng, Alex AceroAbstract:We report our new development of a hidden trajectory model for co-articulated, time-varying patterns of Speech. The model uses bi-directional filtering of vocal tract resonance targets to jointly represent contextual variation and phonetic reduction in Speech Acoustics. A novel maximum-likelihood-based learning algorithm is presented that accurately estimates the distributional parameters of the resonance targets. The results of the estimates are analyzed and shown to be consistent with all the relevant acoustic-phonetic facts and intuitions. Phonetic recognition experiments demonstrate that the model with more rigorous target training outperforms the most recent earlier version of the model, producing 17.5% fewer errors in N-best rescoring.
-
a structured Speech model with continuous hidden dynamics and prediction residual training for tracking vocal tract resonances
International Conference on Acoustics Speech and Signal Processing, 2004Co-Authors: Li Deng, Hagai Attias, Alejandro AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent Speech. The tracking algorithm is based on a version of the structured Speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the prediction residuals. An iterative tracking algorithm is described and evaluated that embeds both the prediction-residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard Speech data demonstrate high accuracy in the results, as well as the effectiveness of residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics.
-
switching dynamic system models for Speech articulation and Acoustics
2004Co-Authors: Li DengAbstract:A statistical generative model for the Speech process is described that embeds a substantially richer structure than the HMM currently in predominant use for automatic Speech recognition. This switching dynamic-system model generalizes and integrates the HMM and the piece-wise stationary nonlinear dynamic system (state- space) model. Depending on the level and the nature of the switching in the model design, various key properties of the Speech dynamics can be naturally represented in the model. Such properties include the temporal structure of the Speech Acoustics, its causal articulatory movements, and the control of such movements by the multidimensional targets correlated with the phonological (symbolic) units of Speech in terms of overlapping articulatory features.
-
an overlapping feature based phonological model incorporating linguistic constraints applications to Speech recognition
Journal of the Acoustical Society of America, 2002Co-Authors: Li DengAbstract:Modeling phonological units of Speech is a critical issue in Speech recognition. In this paper, our recent development of an overlapping-feature-based phonological model that represents long-span contextual dependency in Speech Acoustics is reported. In this model, high-level linguistic constraints are incorporated in automatic construction of the patterns of feature-overlapping and of the hidden Markov model (HMM) states induced by such patterns. The main linguistic information explored includes word and phrase boundaries, morpheme, syllable, syllable constituent categories, and word stress. A consistent computational framework developed for the construction of the feature-based model and the major components of the model are described. Experimental results on the use of the overlapping-feature model in an HMM-based system for Speech recognition show improvements over the conventional triphone-based phonological model.
Alejandro Acero - One of the best experts on this subject based on the ideXlab platform.
-
adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
IEEE Transactions on Audio Speech and Language Processing, 2007Co-Authors: Li Deng, Hagai Attias, Leo J Lee, Alejandro AceroAbstract:A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent Speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonlinear function is adaptively linearized, and a residual or bias term, which is adaptively trained, is added to the nonlinear function to represent the iteratively reduced piecewise linear approximation error. Details of the piecewise linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard Speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics
-
a structured Speech model with continuous hidden dynamics and prediction residual training for tracking vocal tract resonances
International Conference on Acoustics Speech and Signal Processing, 2004Co-Authors: Li Deng, Hagai Attias, Alejandro AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent Speech. The tracking algorithm is based on a version of the structured Speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the prediction residuals. An iterative tracking algorithm is described and evaluated that embeds both the prediction-residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard Speech data demonstrate high accuracy in the results, as well as the effectiveness of residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics.
Wim T J L Pouw - One of the best experts on this subject based on the ideXlab platform.
-
the quantification of gesture Speech synchrony a tutorial and validation of multimodal data acquisition using device based and video based motion tracking
Behavior Research Methods, 2020Co-Authors: Wim T J L Pouw, James P Trujillo, James A DixonAbstract:There is increasing evidence that hand gestures and Speech synchronize their activity on multiple dimensions and timescales. For example, gesture's kinematic peaks (e.g., maximum speed) are coupled with prosodic markers in Speech. Such coupling operates on very short timescales at the level of syllables (200 ms), and therefore requires high-resolution measurement of gesture kinematics and Speech Acoustics. High-resolution Speech analysis is common for gesture studies, given that field's classic ties with (psycho)linguistics. However, the field has lagged behind in the objective study of gesture kinematics (e.g., as compared to research on instrumental action). Often kinematic peaks in gesture are measured by eye, where a "moment of maximum effort" is determined by several raters. In the present article, we provide a tutorial on more efficient methods to quantify the temporal properties of gesture kinematics, in which we focus on common challenges and possible solutions that come with the complexities of studying multimodal language. We further introduce and compare, using an actual gesture dataset (392 gesture events), the performance of two video-based motion-tracking methods (deep learning vs. pixel change) against a high-performance wired motion-tracking system (Polhemus Liberty). We show that the videography methods perform well in the temporal estimation of kinematic peaks, and thus provide a cheap alternative to expensive motion-tracking systems. We hope that the present article incites gesture researchers to embark on the widespread objective study of gesture kinematics and their relation to Speech.
James A Dixon - One of the best experts on this subject based on the ideXlab platform.
-
the quantification of gesture Speech synchrony a tutorial and validation of multimodal data acquisition using device based and video based motion tracking
Behavior Research Methods, 2020Co-Authors: Wim T J L Pouw, James P Trujillo, James A DixonAbstract:There is increasing evidence that hand gestures and Speech synchronize their activity on multiple dimensions and timescales. For example, gesture's kinematic peaks (e.g., maximum speed) are coupled with prosodic markers in Speech. Such coupling operates on very short timescales at the level of syllables (200 ms), and therefore requires high-resolution measurement of gesture kinematics and Speech Acoustics. High-resolution Speech analysis is common for gesture studies, given that field's classic ties with (psycho)linguistics. However, the field has lagged behind in the objective study of gesture kinematics (e.g., as compared to research on instrumental action). Often kinematic peaks in gesture are measured by eye, where a "moment of maximum effort" is determined by several raters. In the present article, we provide a tutorial on more efficient methods to quantify the temporal properties of gesture kinematics, in which we focus on common challenges and possible solutions that come with the complexities of studying multimodal language. We further introduce and compare, using an actual gesture dataset (392 gesture events), the performance of two video-based motion-tracking methods (deep learning vs. pixel change) against a high-performance wired motion-tracking system (Polhemus Liberty). We show that the videography methods perform well in the temporal estimation of kinematic peaks, and thus provide a cheap alternative to expensive motion-tracking systems. We hope that the present article incites gesture researchers to embark on the widespread objective study of gesture kinematics and their relation to Speech.
Hagai Attias - One of the best experts on this subject based on the ideXlab platform.
-
adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
IEEE Transactions on Audio Speech and Language Processing, 2007Co-Authors: Li Deng, Hagai Attias, Leo J Lee, Alejandro AceroAbstract:A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent Speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonlinear function is adaptively linearized, and a residual or bias term, which is adaptively trained, is added to the nonlinear function to represent the iteratively reduced piecewise linear approximation error. Details of the piecewise linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard Speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics
-
a structured Speech model with continuous hidden dynamics and prediction residual training for tracking vocal tract resonances
International Conference on Acoustics Speech and Signal Processing, 2004Co-Authors: Li Deng, Hagai Attias, Alejandro AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent Speech. The tracking algorithm is based on a version of the structured Speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the prediction residuals. An iterative tracking algorithm is described and evaluated that embeds both the prediction-residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard Speech data demonstrate high accuracy in the results, as well as the effectiveness of residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in Speech Acoustics.