The Experts below are selected from a list of 65139 Experts worldwide ranked by ideXlab platform

Kuldip K Paliwal - One of the best experts on this subject based on the ideXlab platform.

  • on training targets for deep learning approaches to clean speech Magnitude Spectrum estimation
    Journal of the Acoustical Society of America, 2021
    Co-Authors: Aaron Nicolson, Kuldip K Paliwal
    Abstract:

    Estimation of the clean speech short-time Magnitude Spectrum (MS) is key for speech enhancement and separation. Moreover, an automatic speech recognition (ASR) system that employs a front-end relies on clean speech MS estimation to remain robust. Training targets for deep learning approaches to clean speech MS estimation fall into three categories: computational auditory scene analysis (CASA), MS, and minimum mean square error (MMSE) estimator training targets. The choice of the training target can have a significant impact on speech enhancement/separation and robust ASR performance. Motivated by this, the training target that produces enhanced/separated speech at the highest quality and intelligibility and that which is best for an ASR front-end is found. Three different deep neural network (DNN) types and two datasets, which include real-world nonstationary and coloured noise sources at multiple signal-to-noise ratio (SNR) levels, were used for evaluation. Ten objective measures were employed, including the word error rate of the Deep Speech ASR system. It is found that training targets that estimate the a priori SNR for MMSE estimators produce the highest objective quality scores. Moreover, it is established that the gain of MMSE estimators and the ideal amplitude mask produce the highest objective intelligibility scores and are most suitable for an ASR front-end.

  • the importance of phase in speech enhancement
    Speech Communication, 2011
    Co-Authors: Kuldip K Paliwal, Kamil Wojcicki, Benjamin Shannon
    Abstract:

    Typical speech enhancement methods, based on the short-time Fourier analysis-modification-synthesis (AMS) framework, modify only the Magnitude Spectrum and keep the phase Spectrum unchanged. In this paper our aim is to show that by modifying the phase Spectrum in the enhancement process the quality of the resulting speech can be improved. For this we use analysis windows of 32ms duration and investigate a number of approaches to phase Spectrum computation. These include the use of matched or mismatched analysis windows for Magnitude and phase spectra estimation during AMS processing, as well as the phase Spectrum compensation (PSC) method. We consider four cases and conduct a series of objective and subjective experiments that examine the importance of the phase Spectrum for speech quality in a systematic manner. In the first (oracle) case, our goal is to determine maximum speech quality improvements achievable when accurate phase Spectrum estimates are available, but when no enhancement is performed on the Magnitude Spectrum. For this purpose speech stimuli are constructed, where (during AMS processing) the phase Spectrum is computed from clean speech, while the Magnitude Spectrum is computed from noisy speech. While such a situation does not arise in practice, it does provide us with a useful insight into how much a precise knowledge of the phase Spectrum can contribute towards speech quality. In this first case, matched and mismatched analysis window approaches are investigated. Particular attention is given to the choice of analysis window type used during phase Spectrum computation, where the effect of spectral dynamic range on speech quality is examined. In the second (non-oracle) case, we consider a more realistic scenario where only the noisy spectra (observable in practice) is available. We study the potential of the mismatched window approach for speech quality improvements in this non-oracle case. We would also like to determine how much room for improvement exists between this case and the best (oracle) case. In the third case, we use the PSC algorithm to enhance the phase Spectrum. We compare this approach with the oracle and non-oracle matched and mismatched window techniques investigated in the preceding cases. While in the first three cases we consider the usefulness of various approaches to phase Spectrum computation within the AMS framework when noisy Magnitude Spectrum is used, in the fourth case we examine the usefulness of these techniques when enhanced Magnitude Spectrum is employed. Our aim (in the context of traditional Magnitude Spectrum-based enhancement methods) is to determine how much benefit in terms of speech quality can be attained by also processing the phase Spectrum. For this purpose, the minimum mean-square error (MMSE) short-time spectral amplitude (STSA) estimates are employed instead of noisy Magnitude spectra. The results of the oracle experiments show that accurate phase Spectrum estimates can considerably contribute towards speech quality, as well as that the use of mismatched analysis windows (in the computation of the Magnitude and phase spectra) provides significant improvements in both objective and subjective speech quality - especially, when the choice of analysis window used for phase Spectrum computation is carefully considered. The mismatched window approach was also found to improve speech quality in the non-oracle case. While the improvements were found to be statistically significant, they were only modest compared to those observed in the oracle case. This suggests that research into better phase Spectrum estimation algorithms, while a challenging task, could be worthwhile. The results of the PSC experiments indicate that the PSC method achieves better speech quality improvements than the other non-oracle methods considered. The results of the MMSE experiments suggest that accurate phase Spectrum estimates have a potential to significantly improve performance of existing Magnitude Spectrum-based methods. Out of the non-oracle approaches considered, the combination of the MMSE STSA method with the PSC algorithm produced significantly better speech quality improvements than those achieved by these methods individually.

  • Role of modulation Magnitude and phase Spectrum towards speech intelligibility
    Speech Communication, 2011
    Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil Wojcicki
    Abstract:

    In this paper our aim is to investigate the properties of the modulation domain and more specifically, to evaluate the relative contributions of the modulation Magnitude and phase spectra towards speech intelligibility. For this purpose, we extend the traditional (acoustic domain) analysis-modification-synthesis framework to include modulation domain processing. We use this framework to construct stimuli that retain only selected spectral components, for the purpose of objective and subjective intelligibility tests. We conduct three experiments. In the first, we investigate the relative contributions to intelligibility of the modulation Magnitude, modulation phase, and acoustic phase spectra. In the second experiment, the effect of modulation frame duration on intelligibility for processing of the modulation Magnitude Spectrum is investigated. In the third experiment, the effect of modulation frame duration on intelligibility for processing of the modulation phase Spectrum is investigated. Results of these experiments show that both the modulation Magnitude and phase spectra are important for speech intelligibility, and that significant improvement is gained by the inclusion of acoustic phase information. They also show that smaller modulation frame durations improve intelligibility when processing the modulation Magnitude Spectrum, while longer frame durations improve intelligibility when processing the modulation phase Spectrum.

  • single channel speech enhancement using mmse estimation of short time modulation Magnitude Spectrum
    Conference of the International Speech Communication Association, 2011
    Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil Wojcicki
    Abstract:

    In this paper we investigate the enhancement of speech by applying MMSE short-time spectral Magnitude estimation in the modulation domain. For this purpose, the traditional analysismodification-synthesis framework is extended to include modulation domain processing. We compensate the noisy modulation Spectrum for additive noise distortion by applying the MMSE short-time spectral Magnitude estimation algorithm in the modulation domain. Subjective experiments were conducted to compare the quality of stimuli processed by the MMSE modulation Magnitude estimator to those processed using the MMSE acoustic Magnitude estimator and the modulation spectral subtraction method. The proposed method is shown to have better noise suppression than MMSE acoustic Magnitude estimation, and improved speech quality compared to modulation domain spectral subtraction. Index Terms: speech enhancement, MMSE short-time spectral Magnitude estimator (AME), modulation Spectrum, MMSE shorttime modulation Magnitude estimator (MME), modulation domain, analysis-modification-synthesis (AMS)

  • product of power Spectrum and group delay function for speech recognition
    International Conference on Acoustics Speech and Signal Processing, 2004
    Co-Authors: Donglai Zhu, Kuldip K Paliwal
    Abstract:

    Mel-frequency cepstral coefficients (MFCCs) are the most widely used features for speech recognition. These are derived from the power Spectrum of the speech signal. Recently, the cepstral features derived from the modified group delay function (MGDF) have been studied by Murthy and Gadde (Proc. ICASSP, vol.1, p.68-71, 2003) for speech recognition. In this paper, we propose to use the product of the power Spectrum and the group delay function (GDF), and derive the MFCCs from the product Spectrum. This Spectrum combines the information from the Magnitude Spectrum as well as the phase Spectrum. The MFCCs of the MGDF are also investigated in this paper. Results show that the cepstral features derived from the power Spectrum perform better than that from the MGDF, and the product Spectrum based features provide the best performance.

Hema A Murthy - One of the best experts on this subject based on the ideXlab platform.

  • an analysis of the high resolution property of group delay function with applications to audio signal processing
    Speech Communication, 2016
    Co-Authors: Jilt Sebastian, Manoj P Kumar, Hema A Murthy
    Abstract:

    Theoretical proof for the high resolution property of group delay function.n dB bandwidth of group delay function is lesser than that of Magnitude Spectrum.Extend the property for multi-resonator systems using empirical measures.Group delay is better compared to Magnitude Spectrum on three applications. This paper provides a new insight into the high resolution property of the negative derivative of the phase response of a system. Group delay functions have been proposed and applied successfully as an alternative to conventional Magnitude Spectrum based applications in speech and music processing. One of the reasons claimed for its superior performance is the high spectral resolution. Most of the existing work use empirical analysis to show this property. In this paper, we show mathematically that for a single resonator, the ratio of the value of the peak in the Magnitude Spectrum to the value at a frequency that is n dB below the peak, is always much lower than the ratio of that of the minimum phase group delay Spectrum. The results are extended for multiple resonators using numerical analyses. The theoretical results are reinforced using three applications, namely, pitch estimation, formant estimation and onset detection. The average deviation from the location of the pitch value/formant value/musical onset is about 53% lower than that of similar techniques that use the Magnitude Spectrum of the signal.

  • incorporating acoustic feature diversity into the linguistic search space for syllable based speech recognition
    European Signal Processing Conference, 2008
    Co-Authors: R Ramya, Rajesh M Hegde, Hema A Murthy
    Abstract:

    Acoustic features derived from the short time Magnitude and phase Spectrum provide complementary information. In this paper, we discuss the significance of incorporating this diverse information into the linguistic search space for syllable based speech recognition. The diversity of group delay acoustic features computed from the phase Spectrum, and MFCC computed from the Magnitude Spectrum, is first illustrated in a lower dimensional feature space. Motivated by this diversity of information in the acoustic feature space, we derive syllable-feature pairs. The selection of syllable-feature pairs is based on isolated syllable recognition results, computed apriori using the two acoustic feature streams. During the recognition process, based on the syllable-feature pair information likelihoods are appropriately weighted using a weighted likelihood scheme. The syllable lattice is now rescored using these weighted syllable-feature pairs in the linguistic search space. This technique of appropriately weighting the relevant acoustic feature for each syllable during the decoding process in the linguistic search space, yields reduced word error rate (WER), for experiments conducted on the TIMIT and the DBIL databases.

  • automatic segmentation of continuous speech using minimum phase group delay functions
    Speech Communication, 2004
    Co-Authors: Kamakshi V Prasad, T Nagarajan, Hema A Murthy
    Abstract:

    Abstract In this paper, we present a new algorithm to automatically segment a continuous speech signal into syllable-like segments. The algorithm for segmentation is based on processing the short-term energy function of the continuous speech signal. The short-term energy function is a positive function and can therefore be processed in a manner similar to that of the Magnitude Spectrum. In this paper, we employ an algorithm, based on group delay processing of the Magnitude Spectrum to determine segment boundaries in the speech signal. The experiments have been carried out on TIMIT and TIDIGITS databases. The error in segment boundary is ⩽20% of syllable duration for 70% of the syllables. In addition to true segments, an overall 5% insertions and deletions have also been observed.

  • segmentation of speech into syllable like units
    Conference of the International Speech Communication Association, 2003
    Co-Authors: T Nagarajan, Hema A Murthy, Rajesh M Hegde
    Abstract:

    In the development of a syllable-centric ASR system, segmentation of the acoustic signal into syllabic units is an important stage. This paper presents a minimum phase group delay based approach to segment spontaneous speech into syllablelike units. Here, three different minimum phase signals are derived from the short term energy functions of three sub-bands of speech signals, as if it were a Magnitude Spectrum. The experiments are carried out on Switchboard and OGI-MLTS corpus and the error in segmentation is found to be utmost 40msec for 85% of the syllable segments.

J F Sturm - One of the best experts on this subject based on the ideXlab platform.

  • linear matrix inequality formulation of spectral mask constraints with applications to fir filter design
    IEEE Transactions on Signal Processing, 2002
    Co-Authors: T N Davidson, Zhiquan Luo, J F Sturm
    Abstract:

    The design of a finite impulse response (FIR) filter often involves a spectral "mask" that the Magnitude Spectrum must satisfy. The mask specifies upper and lower bounds at each frequency and, hence, yields an infinite number of constraints. In current practice, spectral masks are often approximated by discretization, but in this paper, we derive a result that allows us to precisely enforce piecewise constant and piecewise trigonometric polynomial masks in a finite and convex manner via linear matrix inequalities. While this result is theoretically satisfying in that it allows us to avoid the heuristic approximations involved in discretization techniques, it is also of practical interest because it generates competitive design algorithms (based on interior point methods) for a diverse class of FIR filtering and narrowband beamforming problems. The examples we provide include the design of standard linear and nonlinear phase FIR filters, robust "chip" waveforms for wireless communications, and narrowband beamformers for linear antenna arrays. Our main result also provides a contribution to system theory, as it is an extension of the well-known positive-real and bounded-real lemmas.

  • linear matrix inequality formulation of spectral mask constraints
    International Conference on Acoustics Speech and Signal Processing, 2001
    Co-Authors: T N Davidson, Zhiquan Luo, J F Sturm
    Abstract:

    The design of a finite impulse response filter often involves a spectral 'mask' which the Magnitude Spectrum must satisfy. This constraint can be awkward because it yields an infinite number of inequality constraints (two for each frequency point). In current practice, spectral masks are often approximated by discretization, but we show that piecewise constant masks can be precisely enforced in a finite and convex manner via linear matrix inequalities. This facilitates the formulation of a diverse class of filter and beamformer design problems as semidefinite programmes. These optimization problems can be efficiently solved using recently developed interior point methods. Our results can be considered as extensions to the well-known positive-real and bounded-real lemmas from the systems and control literature.

Kamil Wojcicki - One of the best experts on this subject based on the ideXlab platform.

  • the importance of phase in speech enhancement
    Speech Communication, 2011
    Co-Authors: Kuldip K Paliwal, Kamil Wojcicki, Benjamin Shannon
    Abstract:

    Typical speech enhancement methods, based on the short-time Fourier analysis-modification-synthesis (AMS) framework, modify only the Magnitude Spectrum and keep the phase Spectrum unchanged. In this paper our aim is to show that by modifying the phase Spectrum in the enhancement process the quality of the resulting speech can be improved. For this we use analysis windows of 32ms duration and investigate a number of approaches to phase Spectrum computation. These include the use of matched or mismatched analysis windows for Magnitude and phase spectra estimation during AMS processing, as well as the phase Spectrum compensation (PSC) method. We consider four cases and conduct a series of objective and subjective experiments that examine the importance of the phase Spectrum for speech quality in a systematic manner. In the first (oracle) case, our goal is to determine maximum speech quality improvements achievable when accurate phase Spectrum estimates are available, but when no enhancement is performed on the Magnitude Spectrum. For this purpose speech stimuli are constructed, where (during AMS processing) the phase Spectrum is computed from clean speech, while the Magnitude Spectrum is computed from noisy speech. While such a situation does not arise in practice, it does provide us with a useful insight into how much a precise knowledge of the phase Spectrum can contribute towards speech quality. In this first case, matched and mismatched analysis window approaches are investigated. Particular attention is given to the choice of analysis window type used during phase Spectrum computation, where the effect of spectral dynamic range on speech quality is examined. In the second (non-oracle) case, we consider a more realistic scenario where only the noisy spectra (observable in practice) is available. We study the potential of the mismatched window approach for speech quality improvements in this non-oracle case. We would also like to determine how much room for improvement exists between this case and the best (oracle) case. In the third case, we use the PSC algorithm to enhance the phase Spectrum. We compare this approach with the oracle and non-oracle matched and mismatched window techniques investigated in the preceding cases. While in the first three cases we consider the usefulness of various approaches to phase Spectrum computation within the AMS framework when noisy Magnitude Spectrum is used, in the fourth case we examine the usefulness of these techniques when enhanced Magnitude Spectrum is employed. Our aim (in the context of traditional Magnitude Spectrum-based enhancement methods) is to determine how much benefit in terms of speech quality can be attained by also processing the phase Spectrum. For this purpose, the minimum mean-square error (MMSE) short-time spectral amplitude (STSA) estimates are employed instead of noisy Magnitude spectra. The results of the oracle experiments show that accurate phase Spectrum estimates can considerably contribute towards speech quality, as well as that the use of mismatched analysis windows (in the computation of the Magnitude and phase spectra) provides significant improvements in both objective and subjective speech quality - especially, when the choice of analysis window used for phase Spectrum computation is carefully considered. The mismatched window approach was also found to improve speech quality in the non-oracle case. While the improvements were found to be statistically significant, they were only modest compared to those observed in the oracle case. This suggests that research into better phase Spectrum estimation algorithms, while a challenging task, could be worthwhile. The results of the PSC experiments indicate that the PSC method achieves better speech quality improvements than the other non-oracle methods considered. The results of the MMSE experiments suggest that accurate phase Spectrum estimates have a potential to significantly improve performance of existing Magnitude Spectrum-based methods. Out of the non-oracle approaches considered, the combination of the MMSE STSA method with the PSC algorithm produced significantly better speech quality improvements than those achieved by these methods individually.

  • Role of modulation Magnitude and phase Spectrum towards speech intelligibility
    Speech Communication, 2011
    Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil Wojcicki
    Abstract:

    In this paper our aim is to investigate the properties of the modulation domain and more specifically, to evaluate the relative contributions of the modulation Magnitude and phase spectra towards speech intelligibility. For this purpose, we extend the traditional (acoustic domain) analysis-modification-synthesis framework to include modulation domain processing. We use this framework to construct stimuli that retain only selected spectral components, for the purpose of objective and subjective intelligibility tests. We conduct three experiments. In the first, we investigate the relative contributions to intelligibility of the modulation Magnitude, modulation phase, and acoustic phase spectra. In the second experiment, the effect of modulation frame duration on intelligibility for processing of the modulation Magnitude Spectrum is investigated. In the third experiment, the effect of modulation frame duration on intelligibility for processing of the modulation phase Spectrum is investigated. Results of these experiments show that both the modulation Magnitude and phase spectra are important for speech intelligibility, and that significant improvement is gained by the inclusion of acoustic phase information. They also show that smaller modulation frame durations improve intelligibility when processing the modulation Magnitude Spectrum, while longer frame durations improve intelligibility when processing the modulation phase Spectrum.

  • single channel speech enhancement using mmse estimation of short time modulation Magnitude Spectrum
    Conference of the International Speech Communication Association, 2011
    Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil Wojcicki
    Abstract:

    In this paper we investigate the enhancement of speech by applying MMSE short-time spectral Magnitude estimation in the modulation domain. For this purpose, the traditional analysismodification-synthesis framework is extended to include modulation domain processing. We compensate the noisy modulation Spectrum for additive noise distortion by applying the MMSE short-time spectral Magnitude estimation algorithm in the modulation domain. Subjective experiments were conducted to compare the quality of stimuli processed by the MMSE modulation Magnitude estimator to those processed using the MMSE acoustic Magnitude estimator and the modulation spectral subtraction method. The proposed method is shown to have better noise suppression than MMSE acoustic Magnitude estimation, and improved speech quality compared to modulation domain spectral subtraction. Index Terms: speech enhancement, MMSE short-time spectral Magnitude estimator (AME), modulation Spectrum, MMSE shorttime modulation Magnitude estimator (MME), modulation domain, analysis-modification-synthesis (AMS)

T N Davidson - One of the best experts on this subject based on the ideXlab platform.

  • linear matrix inequality formulation of spectral mask constraints with applications to fir filter design
    IEEE Transactions on Signal Processing, 2002
    Co-Authors: T N Davidson, Zhiquan Luo, J F Sturm
    Abstract:

    The design of a finite impulse response (FIR) filter often involves a spectral "mask" that the Magnitude Spectrum must satisfy. The mask specifies upper and lower bounds at each frequency and, hence, yields an infinite number of constraints. In current practice, spectral masks are often approximated by discretization, but in this paper, we derive a result that allows us to precisely enforce piecewise constant and piecewise trigonometric polynomial masks in a finite and convex manner via linear matrix inequalities. While this result is theoretically satisfying in that it allows us to avoid the heuristic approximations involved in discretization techniques, it is also of practical interest because it generates competitive design algorithms (based on interior point methods) for a diverse class of FIR filtering and narrowband beamforming problems. The examples we provide include the design of standard linear and nonlinear phase FIR filters, robust "chip" waveforms for wireless communications, and narrowband beamformers for linear antenna arrays. Our main result also provides a contribution to system theory, as it is an extension of the well-known positive-real and bounded-real lemmas.

  • linear matrix inequality formulation of spectral mask constraints
    International Conference on Acoustics Speech and Signal Processing, 2001
    Co-Authors: T N Davidson, Zhiquan Luo, J F Sturm
    Abstract:

    The design of a finite impulse response filter often involves a spectral 'mask' which the Magnitude Spectrum must satisfy. This constraint can be awkward because it yields an infinite number of inequality constraints (two for each frequency point). In current practice, spectral masks are often approximated by discretization, but we show that piecewise constant masks can be precisely enforced in a finite and convex manner via linear matrix inequalities. This facilitates the formulation of a diverse class of filter and beamformer design problems as semidefinite programmes. These optimization problems can be efficiently solved using recently developed interior point methods. Our results can be considered as extensions to the well-known positive-real and bounded-real lemmas from the systems and control literature.