The Experts below are selected from a list of 28260 Experts worldwide ranked by ideXlab platform
B. Yegnanarayana - One of the best experts on this subject based on the ideXlab platform.
-
INTERSPEECH - Acoustic Characteristics of Ejectives in Amharic
2009Co-Authors: Hussien Seid Worku, S. Rajendran, B. YegnanarayanaAbstract:In this paper, a preliminary investigation of the acoustic characteristics of Amharic ejectives in comparison with their unvoiced conjugates is presented. The normalized error from linear Prediction Residual and a zero frequency resonator output are used to locate the instant of release of the oral closure and the instant of the start of voicing, respectively. Amharic ejectives are found to have longer closure duration and smaller VOT than their unvoiced conjugates. Cross-linguistic comparisons reveal that no ejectives of two languages behave acoustically in a similar manner despite similarity in their articulation. Index Terms: Amharic, ejectives, glottalized stops
-
INTERSPEECH - Voice activity detection in degraded speech using excitation source information.
2007Co-Authors: K. Sri Rama Murty, B. Yegnanarayana, S. GuruprasadAbstract:This paper proposes a method for detection of voiced regions from speech signals collected in noisy environment. The proposed method is based on the characteristics of excitation source of speech production. The degraded speech signal is processed by linear Prediction analysis for deriving the linear Prediction Residual. Hilbert envelope of the linear Prediction Residual is processed using covariance analysis to obtain coherentlyadded covariance signal. The periodicity property of the coherently added covariance signal is exploited to detect the voiced regions using autocorrelation analysis. The performance of the proposed voice activity detection algorithm is evaluated under different noise environments and at different levels of degradation.
-
Extraction of speaker-specific excitation information from linear Prediction Residual of speech
Speech Communication, 2006Co-Authors: S. R. Mahadeva Prasanna, Cheedella S. Gupta, B. YegnanarayanaAbstract:In this paper, through different experimental studies we demonstrate that the excitation component of speech can be exploited for speaker recognition studies. Linear Prediction (LP) Residual is used as a representation of excitation information in speech. The speaker-specific information in the excitation of voiced speech is captured using the AutoAssociative Neural Network (AANN) models. The decrease in the error during training and recognizing correct speakers during testing demonstrates that the excitation component of speech contains speaker-specific information and is indeed being captured by the AANN models. The study on the effect of different LP orders demonstrates that for a speech signal sampled at 8 kHz, the LP Residual extracted using LP order in the range 8-20 best represents the speaker-specific excitation information. It is also demonstrated that the proposed speaker recognition system using excitation information and AANN models requires significantly less amount of data both during training as well as testing, compared to the speaker recognition system using vocal tract information. Finally the speaker recognition studies on NIST 2002 database demonstrates that even though, the recognition performance from the excitation information alone is poor, when combined with evidence from vocal tract information, there is significant improvement in the performance. This result demonstrates the complementary nature of the excitation component of speech.
-
Enhancement of reverberant speech using LP Residual signal
IEEE Transactions on Speech and Audio Processing, 2000Co-Authors: B. Yegnanarayana, P.s. MurthyAbstract:We propose a new method of processing speech degraded by reverberation. The method is based on analysis of short (2 ms) segments of data to enhance the regions in the speech signal having a high signal-to-reverberant component ratio (SRR). The short segment analysis shows that SRR is different in different segments of speech. The processing method involves identifying and manipulating the linear Prediction Residual signal in three different regions of the speech signal, namely, high SRR region, low SRR region, and only reverberation component region. A weight function is derived to modify the linear Prediction Residual signal. The weighted Residual signal samples are used to excite a time-varying all-pole filter to obtain perceptually enhanced speech. The method is robust to noise present in the recorded speech signal. The performance is illustrated through spectrograms, subjective and objective evaluations.
-
Speech enhancement using linear Prediction Residual
Speech Communication, 1999Co-Authors: B. Yegnanarayana, Carlos Avendano, Hynek Hermansky, P.s. MurthyAbstract:Abstract In this paper we propose a method for enhancement of speech in the presence of additive noise. The objective is to selectively enhance the high signal-to-noise ratio (SNR) regions in the noisy speech in the temporal and spectral domains, without causing significant distortion in the resulting enhanced speech. This is proposed to be done at three different levels. (a) At the gross level, by identifying the regions of speech and noise in the temporal domain. (b) At the finer level, by identifying the regions of high and low SNR portions in the noisy speech. (c) At the short-time spectrum level, by enhancing the spectral peaks over spectral valleys. The basis for the proposed approach is to analyze linear Prediction (LP) Residual signal in short (1–2 ms) segments to determine whether a segment belongs to a noise region or speech region. In the speech regions the inverse spectral flatness factor is significantly higher than in the noisy regions. The LP Residual signal enables us to deal with short segments of data due to uncorrelatedness of the samples. Processing of noisy speech for enhancement involves mostly weighting the LP Residual signal samples. The weighted Residual signal samples are used to excite the time-varying all-pole filter to produce enhanced speech. As the additive noise level in the speech signal is increased, the quality of the resulting enhanced speech decreases progressively due to loss of speech information in the low SNR, high noise regions. Thus the degradation in performance of enhancement is graceful as the overall SNR of the noisy speech is decreased.
Li Deng - One of the best experts on this subject based on the ideXlab platform.
-
tracking vocal tract resonances using a quantized nonlinear function embedded in a temporal constraint
IEEE Transactions on Audio Speech and Language Processing, 2006Co-Authors: Li Deng, Alex Acero, I BazziAbstract:This paper presents a new technique for high-accuracy tracking of vocal-tract resonances (which coincide with formants for nonnasalized vowels) in natural speech. The technique is based on a discretized nonlinear Prediction function, which is embedded in a temporal constraint on the quantized input values over adjacent time frames as the prior knowledge for their temporal behavior. The nonlinear Prediction is constructed, based on its analytical form derived in detail in this paper, as a parameter-free, discrete mapping function that approximates the “forward” relationship from the resonance frequencies and bandwidths to the Linear Predictive Coding (LPC) cepstra of real speech. Discretization of the function permits the “inversion” of the function via a search operation. We further introduce the nonlinear-Prediction Residual, characterized by a multivariate Gaussian vector with trainable mean vectors and covariance matrices, to account for the errors due to the functional approximation. We develop and describe an expectation–maximization (EM)-based algorithm for training the parameters of the Residual, and a dynamic programming-based algorithm for resonance tracking. Details of the algorithm implementation for computation speedup are provided. Experimental results are presented which demonstrate the effectiveness of our new paradigm for tracking vocal-tract resonances. In particular, we show the effectiveness of training the Prediction-Residual parameters in obtaining high-accuracy resonance estimates, especially during consonantal closure.
-
a structured speech model with continuous hidden dynamics and Prediction Residual training for tracking vocal tract resonances
International Conference on Acoustics Speech and Signal Processing, 2004Co-Authors: Li Deng, Hagai Attias, Alejandro AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent speech. The tracking algorithm is based on a version of the structured speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized Prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the Prediction Residuals. An iterative tracking algorithm is described and evaluated that embeds both the Prediction-Residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard speech data demonstrate high accuracy in the results, as well as the effectiveness of Residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics.
-
tracking vocal tract resonances using an analytical nonlinear predictor and a target guided temporal constraint
Conference of the International Speech Communication Association, 2003Co-Authors: Li Deng, I Bazzi, Alex AceroAbstract:A technique for high-accuracy tracking of formants or vocal tract resonances is presented in this paper using a novel nonlinear predictor and using a target-directed temporal constraint. The nonlinear predictor is constructed from a parameter-free, discrete mapping function from the formant (frequencies and bandwidths) space to the LPC-cepstral space, with trainable Residuals. We examine in this study the key role of vocal tract resonance targets in the tracking accuracy. Experimental results show that due to the use of the targets, the tracked formants in the consonantal regions (including closures and short pauses) of the speech utterance exhibit the same dynamic properties as for the vocalic regions, and reflect the underlying vocal tract resonances. The results also demonstrate the effectiveness of training the Prediction-Residual parameters and of incorporating the target-based constraint in obtaining high-accuracy formant estimates, especially for non-sonorant portions of speech.
-
ICASSP (1) - A structured speech model with continuous hidden dynamics and Prediction-Residual training for tracking vocal tract resonances
2004 IEEE International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: Li Deng, Hagai Attias, Leo J. Lee, Alex AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent speech. The tracking algorithm is based on a version of the structured speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized Prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the Prediction Residuals. An iterative tracking algorithm is described and evaluated that embeds both the Prediction-Residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard speech data demonstrate high accuracy in the results, as well as the effectiveness of Residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics.
P.s. Murthy - One of the best experts on this subject based on the ideXlab platform.
-
Enhancement of reverberant speech using LP Residual signal
IEEE Transactions on Speech and Audio Processing, 2000Co-Authors: B. Yegnanarayana, P.s. MurthyAbstract:We propose a new method of processing speech degraded by reverberation. The method is based on analysis of short (2 ms) segments of data to enhance the regions in the speech signal having a high signal-to-reverberant component ratio (SRR). The short segment analysis shows that SRR is different in different segments of speech. The processing method involves identifying and manipulating the linear Prediction Residual signal in three different regions of the speech signal, namely, high SRR region, low SRR region, and only reverberation component region. A weight function is derived to modify the linear Prediction Residual signal. The weighted Residual signal samples are used to excite a time-varying all-pole filter to obtain perceptually enhanced speech. The method is robust to noise present in the recorded speech signal. The performance is illustrated through spectrograms, subjective and objective evaluations.
-
Speech enhancement using linear Prediction Residual
Speech Communication, 1999Co-Authors: B. Yegnanarayana, Carlos Avendano, Hynek Hermansky, P.s. MurthyAbstract:Abstract In this paper we propose a method for enhancement of speech in the presence of additive noise. The objective is to selectively enhance the high signal-to-noise ratio (SNR) regions in the noisy speech in the temporal and spectral domains, without causing significant distortion in the resulting enhanced speech. This is proposed to be done at three different levels. (a) At the gross level, by identifying the regions of speech and noise in the temporal domain. (b) At the finer level, by identifying the regions of high and low SNR portions in the noisy speech. (c) At the short-time spectrum level, by enhancing the spectral peaks over spectral valleys. The basis for the proposed approach is to analyze linear Prediction (LP) Residual signal in short (1–2 ms) segments to determine whether a segment belongs to a noise region or speech region. In the speech regions the inverse spectral flatness factor is significantly higher than in the noisy regions. The LP Residual signal enables us to deal with short segments of data due to uncorrelatedness of the samples. Processing of noisy speech for enhancement involves mostly weighting the LP Residual signal samples. The weighted Residual signal samples are used to excite the time-varying all-pole filter to produce enhanced speech. As the additive noise level in the speech signal is increased, the quality of the resulting enhanced speech decreases progressively due to loss of speech information in the low SNR, high noise regions. Thus the degradation in performance of enhancement is graceful as the overall SNR of the noisy speech is decreased.
Alex Acero - One of the best experts on this subject based on the ideXlab platform.
-
tracking vocal tract resonances using a quantized nonlinear function embedded in a temporal constraint
IEEE Transactions on Audio Speech and Language Processing, 2006Co-Authors: Li Deng, Alex Acero, I BazziAbstract:This paper presents a new technique for high-accuracy tracking of vocal-tract resonances (which coincide with formants for nonnasalized vowels) in natural speech. The technique is based on a discretized nonlinear Prediction function, which is embedded in a temporal constraint on the quantized input values over adjacent time frames as the prior knowledge for their temporal behavior. The nonlinear Prediction is constructed, based on its analytical form derived in detail in this paper, as a parameter-free, discrete mapping function that approximates the “forward” relationship from the resonance frequencies and bandwidths to the Linear Predictive Coding (LPC) cepstra of real speech. Discretization of the function permits the “inversion” of the function via a search operation. We further introduce the nonlinear-Prediction Residual, characterized by a multivariate Gaussian vector with trainable mean vectors and covariance matrices, to account for the errors due to the functional approximation. We develop and describe an expectation–maximization (EM)-based algorithm for training the parameters of the Residual, and a dynamic programming-based algorithm for resonance tracking. Details of the algorithm implementation for computation speedup are provided. Experimental results are presented which demonstrate the effectiveness of our new paradigm for tracking vocal-tract resonances. In particular, we show the effectiveness of training the Prediction-Residual parameters in obtaining high-accuracy resonance estimates, especially during consonantal closure.
-
tracking vocal tract resonances using an analytical nonlinear predictor and a target guided temporal constraint
Conference of the International Speech Communication Association, 2003Co-Authors: Li Deng, I Bazzi, Alex AceroAbstract:A technique for high-accuracy tracking of formants or vocal tract resonances is presented in this paper using a novel nonlinear predictor and using a target-directed temporal constraint. The nonlinear predictor is constructed from a parameter-free, discrete mapping function from the formant (frequencies and bandwidths) space to the LPC-cepstral space, with trainable Residuals. We examine in this study the key role of vocal tract resonance targets in the tracking accuracy. Experimental results show that due to the use of the targets, the tracked formants in the consonantal regions (including closures and short pauses) of the speech utterance exhibit the same dynamic properties as for the vocalic regions, and reflect the underlying vocal tract resonances. The results also demonstrate the effectiveness of training the Prediction-Residual parameters and of incorporating the target-based constraint in obtaining high-accuracy formant estimates, especially for non-sonorant portions of speech.
-
ICASSP (1) - A structured speech model with continuous hidden dynamics and Prediction-Residual training for tracking vocal tract resonances
2004 IEEE International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: Li Deng, Hagai Attias, Leo J. Lee, Alex AceroAbstract:A novel approach is developed for efficient and accurate tracking of vocal tract resonances, which are natural frequencies of the resonator from larynx to lips, in fluent speech. The tracking algorithm is based on a version of the structured speech model consisting of continuous-valued hidden dynamics and a piecewise-linearized Prediction function from resonance frequencies and bandwidths to LPC cepstra. We present details of the piecewise linearization design process and an adaptive training technique for the parameters that characterize the Prediction Residuals. An iterative tracking algorithm is described and evaluated that embeds both the Prediction-Residual training and the piecewise linearization design in an adaptive Kalman filtering framework. Experiments on tracking vocal tract resonances in Switchboard speech data demonstrate high accuracy in the results, as well as the effectiveness of Residual training embedded in the algorithm. Our approach differs from traditional formant trackers in that it provides meaningful results even during consonantal closures when the supra-laryngeal source may cause no spectral prominences in speech acoustics.
Felix Carlos Fernandes - One of the best experts on this subject based on the ideXlab platform.
-
Low Latency Secondary Transforms for Intra/Inter Prediction Residual
IEEE transactions on image processing : a publication of the IEEE Signal Processing Society, 2013Co-Authors: Ankur Saxena, Felix Carlos FernandesAbstract:In this paper, we present a transform scheme where a secondary transform is applied after the conventional DCT for intra as well as inter Prediction residues. Our approach is applicable to any block-based video codec that employs transforms along the horizontal and vertical direction separably. The secondary transform is applied to the lower K ( K=4 or 8) frequency coefficients of the output of conventional DCT at block with dimensions 8 and larger. The proposed transform scheme has low complexity as it is applied only to the top-left portion of the DCT output, especially in the context of large blocks such as 32 × 32 where an alternate non-DCT 32 × 32 transform would have a prohibitive implementation hardware cost. The proposed technique is single-pass, and the choice of whether to use the secondary transform is solely based on the Prediction direction for intra residue, and on transform unit location in the Prediction unit for the inter residue. The scheme requires no additional signaling information or R-D search. Our simulation results show that the proposed transform scheme provides significant BD-rate improvement over the conventional DCT-based coding scheme. Finally, we also show how to implement the proposed secondary transforms with low latency in hardware.
-
ICIP - On secondary transforms for Prediction Residual
2012 19th IEEE International Conference on Image Processing, 2012Co-Authors: Ankur Saxena, Felix Carlos FernandesAbstract:In this paper, we present a transform scheme that utilizes a secondary transform after the conventional DCT for intra as well as inter Prediction residues. Our approach is applicable to any block-based video codec that employs transforms along the horizontal and vertical direction separably. The secondary transform is applied to the lower K (K=4 or 8) frequency coefficients of the output of conventional DCT at block with dimensions 8 and higher. The proposed transform scheme has low complexity as it is applied only to the top-left portion of DCT output, especially in the context of large blocks such as 32×32 where an alternate transform of size 32×32 other than DCT would be expensive to be implemented in hardware. The proposed technique works in single-pass, and the choice of whether to use the secondary transform is solely based on the Prediction direction for intra residue, and transform unit location in the Prediction unit for the inter residue. The scheme requires no additional signaling information or R-D search. Our simulation results show that the proposed transform scheme provides significant BD-Rate improvement over the conventional DCT-based coding scheme for video sequences in the ongoing HEVC standardization.
-
ICASSP - On secondary transforms for intra Prediction Residual
2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012Co-Authors: Ankur Saxena, Felix Carlos FernandesAbstract:In this paper, we present a secondary transform that is applied in a codec after the conventional DCT for all the video-coding intra Prediction modes. Our approach is applicable to any block-based intra Prediction scheme that employs transforms along the horizontal and vertical direction separably. The secondary transform is applied to the lower 8×8 frequency coefficients of the output of conventional DCT at block sizes 8×8 and higher. The proposed transform scheme has low complexity as it is applied only to the top-left portion of DCT output, especially in the context of large blocks such as 32×32 where an alternate transform of size 32×32 other than DCT would be expensive to be implemented in hardware. The proposed technique works in single-pass, and the choice of when to use the secondary transform is solely based on the intra Prediction mode and requires no additional signaling information or R-D search. Our simulation results show that the proposed transform scheme provides significant BD-Rate improvement over the conventional DCT-based coding scheme for intra Prediction of video sequences in the ongoing HEVC standardization.