The Experts below are selected from a list of 324 Experts worldwide ranked by ideXlab platform
Kuldip K Paliwal - One of the best experts on this subject based on the ideXlab platform.
-
Role of modulation magnitude and Phase Spectrum towards speech intelligibility
Speech Communication, 2011Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil WojcickiAbstract:In this paper our aim is to investigate the properties of the modulation domain and more specifically, to evaluate the relative contributions of the modulation magnitude and Phase spectra towards speech intelligibility. For this purpose, we extend the traditional (acoustic domain) analysis-modification-synthesis framework to include modulation domain processing. We use this framework to construct stimuli that retain only selected spectral components, for the purpose of objective and subjective intelligibility tests. We conduct three experiments. In the first, we investigate the relative contributions to intelligibility of the modulation magnitude, modulation Phase, and acoustic Phase spectra. In the second experiment, the effect of modulation frame duration on intelligibility for processing of the modulation magnitude Spectrum is investigated. In the third experiment, the effect of modulation frame duration on intelligibility for processing of the modulation Phase Spectrum is investigated. Results of these experiments show that both the modulation magnitude and Phase spectra are important for speech intelligibility, and that significant improvement is gained by the inclusion of acoustic Phase information. They also show that smaller modulation frame durations improve intelligibility when processing the modulation magnitude Spectrum, while longer frame durations improve intelligibility when processing the modulation Phase Spectrum.
-
INTERSPEECH - Noise driven short-time Phase Spectrum compensation procedure for speech enhancement
2008Co-Authors: Anthony Phillip Stark, Kamil Wojcicki, James Lyons, Kuldip K PaliwalAbstract:Typical speech enhancement algorithms operate on the shorttime magnitude Spectrum, while keeping the short-time Phase Spectrum unchanged for synthesis. Recently, a novel approach to speech enhancement has been proposed where the noisy magnitude Spectrum is recombined with a changed Phase Spectrum to produce a modified complex Spectrum. During synthesis the low energy components of the modified complex Spectrum cancel out more than the high energy components, thus reducing background noise. In the present work, a procedure that employs noise estimates to compensate the Phase Spectrum for additive noise distortion is formulated. The proposed approach is objectively evaluated against several popular speech enhancement methods under various noise conditions and is shown to compare favourably
-
Short-time Phase Spectrum in speech processing: A review and some experimental results
Digital Signal Processing, 2007Co-Authors: Leigh Alsteris, Kuldip K PaliwalAbstract:Incorporating information from the short-time Phase Spectrum into a feature set for automatic speech recognition (ASR) may possibly serve to improve recognition accuracy. Currently, however, it is common practice to discard this information in favour of features that are derived purely from the short-time magnitude Spectrum. There are two reasons for this: (1) the results of some well-known human listening experiments have indicated that the short-time Phase Spectrum conveys a negligible amount of intelligibility at the small window durations of 20-40 ms used for ASR spectral analysis, and (2) using the short-time Phase Spectrum directly for ASR has proven difficult from a signal processing viewpoint, due to Phase-wrapping and other problems. In this article, we explore the possibility of using short-time Phase Spectrum information for ASR by considering the two points mentioned above. To address the first point, we review the results of our own set of human listening experiments. Contrary to previous studies, our results indicate that the short-time Phase Spectrum can indeed contribute significantly to speech intelligibility over small window durations of 20-40 ms. Also, the results of these listening experiments, in addition to some ASR experiments, indicate that at least part of this intelligibility may be supplementary to that provided by the short-time magnitude Spectrum. To address the second point (i.e., the signal processing difficulties), we suggest that it may be necessary to transform the short-time Phase Spectrum into a more physically meaningful representation from which useful features could possibly be extracted. Specifically, we investigate the frequency-derivative (or group delay function, GDF) and the time-derivative (or instantaneous frequency distribution, IFD) as potential candidates for this intermediate representation. We review our recent work, where we have performed various experiments which show that the GDF and IFD may be useful for ASR. In our recent work, we have also conducted several ASR experiments to test a feature set derived from the GDF. We found that, in most cases, these features perform worse than the standard MFCC features. Therefore, we suggest that a short-time Phase Spectrum feature set may ultimately be derived from a concatenation of information from both the GDF and IFD representations. For best performance, the feature set may also need to be concatenated with short-time magnitude Spectrum information. Further to addressing the two aforementioned points, we also discuss a number of other speech applications in which the short-time Phase Spectrum has proven to be very useful. We believe that an appreciation for how the short-time Phase Spectrum has been used for other tasks, in addition to the results of our own experiments, will provoke fellow researchers to also investigate its potential for use in ASR.
-
Further intelligibility results from human listening tests using the short-time Phase Spectrum
Speech Communication, 2006Co-Authors: Leigh Alsteris, Kuldip K PaliwalAbstract:State-of-the-art automatic speech recognition systems (ASRs) use only the short-time magnitude Spectrum for feature extraction; the short-time Phase Spectrum is generally ignored in these systems. Results from our recent human listening tests indicate that the short-time Phase Spectrum can significantly contribute to speech intelligibility over small window durations (i.e., 20-40 ms). This is an interesting result, indicating the possible usefulness of the short-time Phase Spectrum for ASR, which commonly employs small window durations of 20-40 ms for spectral analysis. In this paper, we continue our investigation of the short-time Phase Spectrum. We explore the use of partial short-time Phase Spectrum information, in the absence of all the short-time magnitude Spectrum information, for intelligible signal reconstruction. We create two types of stimuli; one in which its frequency-derivative (i.e., group delay function, GDF) is preserved and another in which its time-derivative (i.e., instantaneous frequency distribution, IFD) is preserved. We do this to determine the contribution that each of these derivatives provides toward intelligibility. Reconstructing stimuli from knowledge of only the GDF or only the IFD results in poor intelligibility. However, when we create stimuli using knowledge of both the GDF and the IFD, reasonable intelligibility is obtained. In light of these results, we conclude that both the GDF and IFD components of the short-time Phase Spectrum are needed to reconstruct an intelligible signal. In addition, we also perform some experiments to quantify the intelligibility of stimuli reconstructed from the short-time Phase and magnitude spectra of noisy speech. The intelligibility of stimuli constructed from either the short-time magnitude Spectrum or the short-time Phase Spectrum degrades at a similar rate under increasing noise levels. The intelligibility of the original signals under noisy conditions also degrades with increased noise, but in all cases the intelligibility is superior to that provided by the stimuli constructed from the separate short-time components. Therefore, we argue that knowledge of both short-time magnitude and Phase Spectrum information results in superior human speech recognition performance.
-
On the usefulness of STFT Phase Spectrum in human listening tests
Speech Communication, 2005Co-Authors: Kuldip K Paliwal, Leigh AlsterisAbstract:Abstract The short-time Fourier transform (STFT) of a speech signal has two components: the magnitude Spectrum and the Phase Spectrum. In this paper, the relative importance of short-time magnitude and Phase spectra for speech perception is investigated. Human perception experiments are conducted to measure intelligibility of speech stimuli synthesized either from magnitude spectra or Phase spectra. It is traditionally believed that the magnitude Spectrum plays a dominant role for small window durations (20–40 ms); while the Phase Spectrum is more important for large window durations (>1 s). It is shown in this paper that even for small window durations, the Phase Spectrum can contribute to speech intelligibility as much as the magnitude Spectrum if the analysis–modification–synthesis parameters are properly selected.
Leigh Alsteris - One of the best experts on this subject based on the ideXlab platform.
-
Short-time Phase Spectrum in speech processing: A review and some experimental results
Digital Signal Processing, 2007Co-Authors: Leigh Alsteris, Kuldip K PaliwalAbstract:Incorporating information from the short-time Phase Spectrum into a feature set for automatic speech recognition (ASR) may possibly serve to improve recognition accuracy. Currently, however, it is common practice to discard this information in favour of features that are derived purely from the short-time magnitude Spectrum. There are two reasons for this: (1) the results of some well-known human listening experiments have indicated that the short-time Phase Spectrum conveys a negligible amount of intelligibility at the small window durations of 20-40 ms used for ASR spectral analysis, and (2) using the short-time Phase Spectrum directly for ASR has proven difficult from a signal processing viewpoint, due to Phase-wrapping and other problems. In this article, we explore the possibility of using short-time Phase Spectrum information for ASR by considering the two points mentioned above. To address the first point, we review the results of our own set of human listening experiments. Contrary to previous studies, our results indicate that the short-time Phase Spectrum can indeed contribute significantly to speech intelligibility over small window durations of 20-40 ms. Also, the results of these listening experiments, in addition to some ASR experiments, indicate that at least part of this intelligibility may be supplementary to that provided by the short-time magnitude Spectrum. To address the second point (i.e., the signal processing difficulties), we suggest that it may be necessary to transform the short-time Phase Spectrum into a more physically meaningful representation from which useful features could possibly be extracted. Specifically, we investigate the frequency-derivative (or group delay function, GDF) and the time-derivative (or instantaneous frequency distribution, IFD) as potential candidates for this intermediate representation. We review our recent work, where we have performed various experiments which show that the GDF and IFD may be useful for ASR. In our recent work, we have also conducted several ASR experiments to test a feature set derived from the GDF. We found that, in most cases, these features perform worse than the standard MFCC features. Therefore, we suggest that a short-time Phase Spectrum feature set may ultimately be derived from a concatenation of information from both the GDF and IFD representations. For best performance, the feature set may also need to be concatenated with short-time magnitude Spectrum information. Further to addressing the two aforementioned points, we also discuss a number of other speech applications in which the short-time Phase Spectrum has proven to be very useful. We believe that an appreciation for how the short-time Phase Spectrum has been used for other tasks, in addition to the results of our own experiments, will provoke fellow researchers to also investigate its potential for use in ASR.
-
Further intelligibility results from human listening tests using the short-time Phase Spectrum
Speech Communication, 2006Co-Authors: Leigh Alsteris, Kuldip K PaliwalAbstract:State-of-the-art automatic speech recognition systems (ASRs) use only the short-time magnitude Spectrum for feature extraction; the short-time Phase Spectrum is generally ignored in these systems. Results from our recent human listening tests indicate that the short-time Phase Spectrum can significantly contribute to speech intelligibility over small window durations (i.e., 20-40 ms). This is an interesting result, indicating the possible usefulness of the short-time Phase Spectrum for ASR, which commonly employs small window durations of 20-40 ms for spectral analysis. In this paper, we continue our investigation of the short-time Phase Spectrum. We explore the use of partial short-time Phase Spectrum information, in the absence of all the short-time magnitude Spectrum information, for intelligible signal reconstruction. We create two types of stimuli; one in which its frequency-derivative (i.e., group delay function, GDF) is preserved and another in which its time-derivative (i.e., instantaneous frequency distribution, IFD) is preserved. We do this to determine the contribution that each of these derivatives provides toward intelligibility. Reconstructing stimuli from knowledge of only the GDF or only the IFD results in poor intelligibility. However, when we create stimuli using knowledge of both the GDF and the IFD, reasonable intelligibility is obtained. In light of these results, we conclude that both the GDF and IFD components of the short-time Phase Spectrum are needed to reconstruct an intelligible signal. In addition, we also perform some experiments to quantify the intelligibility of stimuli reconstructed from the short-time Phase and magnitude spectra of noisy speech. The intelligibility of stimuli constructed from either the short-time magnitude Spectrum or the short-time Phase Spectrum degrades at a similar rate under increasing noise levels. The intelligibility of the original signals under noisy conditions also degrades with increased noise, but in all cases the intelligibility is superior to that provided by the stimuli constructed from the separate short-time components. Therefore, we argue that knowledge of both short-time magnitude and Phase Spectrum information results in superior human speech recognition performance.
-
On the usefulness of STFT Phase Spectrum in human listening tests
Speech Communication, 2005Co-Authors: Kuldip K Paliwal, Leigh AlsterisAbstract:Abstract The short-time Fourier transform (STFT) of a speech signal has two components: the magnitude Spectrum and the Phase Spectrum. In this paper, the relative importance of short-time magnitude and Phase spectra for speech perception is investigated. Human perception experiments are conducted to measure intelligibility of speech stimuli synthesized either from magnitude spectra or Phase spectra. It is traditionally believed that the magnitude Spectrum plays a dominant role for small window durations (20–40 ms); while the Phase Spectrum is more important for large window durations (>1 s). It is shown in this paper that even for small window durations, the Phase Spectrum can contribute to speech intelligibility as much as the magnitude Spectrum if the analysis–modification–synthesis parameters are properly selected.
-
usefulness of Phase Spectrum in human speech perception
Conference of the International Speech Communication Association, 2003Co-Authors: Kuldip K Paliwal, Leigh AlsterisAbstract:Short-time Fourier transform of speech signal has two components: magnitude Spectrum and Phase Spectrum. In this paper, relative importance of short-time magnitude and Phase spectra on speech perception is investigated. Human perception experiments are conducted to measure intelligibility of speech tokens synthesized either from magnitude Spectrum or Phase Spectrum. It is traditionally believed that magnitude Spectrum plays a dominant role for shorter windows (20-30 ms); while Phase Spectrum is more important for longer windows (128-3500 ms). It is shown in this paper that even for shorter windows, Phase Spectrum can contribute to speech intelligibility as much as the magnitude Spectrum if the shape of the window function is properly selected.
-
INTERSPEECH - USEFULNESS OF Phase Spectrum IN HUMAN SPEECH PERCEPTION
2003Co-Authors: Kuldip K Paliwal, Leigh AlsterisAbstract:Short-time Fourier transform of speech signal has two components: magnitude Spectrum and Phase Spectrum. In this paper, relative importance of short-time magnitude and Phase spectra on speech perception is investigated. Human perception experiments are conducted to measure intelligibility of speech tokens synthesized either from magnitude Spectrum or Phase Spectrum. It is traditionally believed that magnitude Spectrum plays a dominant role for shorter windows (20-30 ms); while Phase Spectrum is more important for longer windows (128-3500 ms). It is shown in this paper that even for shorter windows, Phase Spectrum can contribute to speech intelligibility as much as the magnitude Spectrum if the shape of the window function is properly selected.
Kamil Wojcicki - One of the best experts on this subject based on the ideXlab platform.
-
Role of modulation magnitude and Phase Spectrum towards speech intelligibility
Speech Communication, 2011Co-Authors: Kuldip K Paliwal, Belinda Schwerin, Kamil WojcickiAbstract:In this paper our aim is to investigate the properties of the modulation domain and more specifically, to evaluate the relative contributions of the modulation magnitude and Phase spectra towards speech intelligibility. For this purpose, we extend the traditional (acoustic domain) analysis-modification-synthesis framework to include modulation domain processing. We use this framework to construct stimuli that retain only selected spectral components, for the purpose of objective and subjective intelligibility tests. We conduct three experiments. In the first, we investigate the relative contributions to intelligibility of the modulation magnitude, modulation Phase, and acoustic Phase spectra. In the second experiment, the effect of modulation frame duration on intelligibility for processing of the modulation magnitude Spectrum is investigated. In the third experiment, the effect of modulation frame duration on intelligibility for processing of the modulation Phase Spectrum is investigated. Results of these experiments show that both the modulation magnitude and Phase spectra are important for speech intelligibility, and that significant improvement is gained by the inclusion of acoustic Phase information. They also show that smaller modulation frame durations improve intelligibility when processing the modulation magnitude Spectrum, while longer frame durations improve intelligibility when processing the modulation Phase Spectrum.
-
INTERSPEECH - Noise driven short-time Phase Spectrum compensation procedure for speech enhancement
2008Co-Authors: Anthony Phillip Stark, Kamil Wojcicki, James Lyons, Kuldip K PaliwalAbstract:Typical speech enhancement algorithms operate on the shorttime magnitude Spectrum, while keeping the short-time Phase Spectrum unchanged for synthesis. Recently, a novel approach to speech enhancement has been proposed where the noisy magnitude Spectrum is recombined with a changed Phase Spectrum to produce a modified complex Spectrum. During synthesis the low energy components of the modified complex Spectrum cancel out more than the high energy components, thus reducing background noise. In the present work, a procedure that employs noise estimates to compensate the Phase Spectrum for additive noise distortion is formulated. The proposed approach is objectively evaluated against several popular speech enhancement methods under various noise conditions and is shown to compare favourably
Hou Zi-qiang - One of the best experts on this subject based on the ideXlab platform.
-
ICASSP - The generalized Phase Spectrum method for time delay estimation
ICASSP '84. IEEE International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: Zhao Zhen, Hou Zi-qiangAbstract:The concept of the Generalized Phase Spectrum (GPS) TDE is put forward. The relation between the GPS TDE and the GCC TDE is derived. A multipath signal model is considered. The method of Amplitude Square (AS) weighting for TDE is proposed and the comparison between the AS TDE and the Phase Data (PD) TDE is made in the multipath environment. The results of theoretical calculation and computer simulation experiment show that the performance of the AS weighting TDE is superior to that of the PD TDE.
Cai Han-pen - One of the best experts on this subject based on the ideXlab platform.
-
Thickness Estimates from Instantaneous Phase Spectrum of Poststack Seismic Data
Natural Gas Geoscience, 2014Co-Authors: Cai Han-penAbstract:Thickness estimates using instantaneous Phase Spectrum of post-stack seismic data was explored to estimate bed thickness that is less than tuning thickness.Objective function for thickness estimates,in this method,was constructed on the basis of the instantaneous Phase Spectrum in the available seismic frequency band,in which magnitude and polarity of seismic reflectivity and the dominant frequency of seismic wavelet were not need to be considered.Synthetic data test showed that on the premise of seismic data without noise,bed thickness of less than and greater than tuning thickness is estimated accurately,and bed thickness estimates using the proposed method is not affected by the frequency bandwidth of seismic data and magnitude and polarity of seismic reflectivity,but thickness estimates is contaminated seriously by noise.Real seismic data examples was evaluated to demonstrate that when SNR of seismic data within the effective frequency band is very high,estimation error employing this method,compared with well logging and interpretation data,is less than 10%,providing the basis for drilling deployment of petroleum reservoir exploration and development.