The Experts below are selected from a list of 29076 Experts worldwide ranked by ideXlab platform

Alejandro Acero - One of the best experts on this subject based on the ideXlab platform.

  • adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Li Deng, Leo J Lee, Hagai Attias, Alejandro Acero
    Abstract:

    A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonLinear function is adaptively Linearized, and a residual or bias term, which is adaptively trained, is added to the nonLinear function to represent the iteratively reduced piecewise Linear Approximation Error. Details of the piecewise Linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise Linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics

Kjersti Engan - One of the best experts on this subject based on the ideXlab platform.

  • design of signal expansions for sparse representation
    International Conference on Acoustics Speech and Signal Processing, 2000
    Co-Authors: S.o. Aase, John Hakon Husoy, Karl Skretting, Kjersti Engan
    Abstract:

    Traditional signal decompositions generate signal expansions using the analysis-synthesis setting: the expansion coefficients are found by taking the inner product of the signal with the corresponding analysis vector. In this paper we try to free ourselves from the analysis-synthesis paradigm by concentrating on the synthesis or reconstruction part of the signal expansion. Ignoring the analysis issue completely, we construct sets of synthesis vectors, denoted waveform dictionaries, for sparse signal representation. The objective is to approximate a training signal using a small number of dictionary vectors. Our algorithm optimize the dictionary vectors with respect to the average non-Linear Approximation Error. Using signals from a Gaussian, autoregressive process with correlation factor 0.95, it is demonstrated that for established signal expansions like the Karhunen-Loeve transform, the lapped orthogonal transform, and the biorthogonal 7/9 wavelet, it is possible to improve the Approximation capabilities by up to 30% by optimizing the expansion vectors.

Li Deng - One of the best experts on this subject based on the ideXlab platform.

  • adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Li Deng, Leo J Lee, Hagai Attias, Alejandro Acero
    Abstract:

    A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonLinear function is adaptively Linearized, and a residual or bias term, which is adaptively trained, is added to the nonLinear function to represent the iteratively reduced piecewise Linear Approximation Error. Details of the piecewise Linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise Linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics

S.o. Aase - One of the best experts on this subject based on the ideXlab platform.

  • design of signal expansions for sparse representation
    International Conference on Acoustics Speech and Signal Processing, 2000
    Co-Authors: S.o. Aase, John Hakon Husoy, Karl Skretting, Kjersti Engan
    Abstract:

    Traditional signal decompositions generate signal expansions using the analysis-synthesis setting: the expansion coefficients are found by taking the inner product of the signal with the corresponding analysis vector. In this paper we try to free ourselves from the analysis-synthesis paradigm by concentrating on the synthesis or reconstruction part of the signal expansion. Ignoring the analysis issue completely, we construct sets of synthesis vectors, denoted waveform dictionaries, for sparse signal representation. The objective is to approximate a training signal using a small number of dictionary vectors. Our algorithm optimize the dictionary vectors with respect to the average non-Linear Approximation Error. Using signals from a Gaussian, autoregressive process with correlation factor 0.95, it is demonstrated that for established signal expansions like the Karhunen-Loeve transform, the lapped orthogonal transform, and the biorthogonal 7/9 wavelet, it is possible to improve the Approximation capabilities by up to 30% by optimizing the expansion vectors.

Leo J Lee - One of the best experts on this subject based on the ideXlab platform.

  • adaptive kalman filtering and smoothing for tracking vocal tract resonances using a continuous valued hidden dynamic model
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Li Deng, Leo J Lee, Hagai Attias, Alejandro Acero
    Abstract:

    A novel Kalman filtering/smoothing algorithm is presented for efficient and accurate estimation of vocal tract resonances or formants, which are natural frequencies and bandwidths of the resonator from larynx to lips, in fluent speech. The algorithm uses a hidden dynamic model, with a state-space formulation, where the resonance frequency and bandwidth values are treated as continuous-valued hidden state variables. The observation equation of the model is constructed by an analytical predictive function from the resonance frequencies and bandwidths to LPC cepstra as the observation vectors. This nonLinear function is adaptively Linearized, and a residual or bias term, which is adaptively trained, is added to the nonLinear function to represent the iteratively reduced piecewise Linear Approximation Error. Details of the piecewise Linearization design process are described. An iterative tracking algorithm is presented, which embeds both the adaptive residual training and piecewise Linearization design in the Kalman filtering/smoothing framework. Experiments on estimating resonances in Switchboard speech data show accurate estimation results. In particular, the effectiveness of the adaptive residual training is demonstrated. Our approach provides a solution to the traditional "hidden formant problem," and produces meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics