The Experts below are selected from a list of 4671 Experts worldwide ranked by ideXlab platform
Junichi Yamagishi - One of the best experts on this subject based on the ideXlab platform.
-
Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
arXiv: Audio and Speech Processing, 2020Co-Authors: Yi Zhao, Cheng-i Lai, Jennifer Williams, Erica Cooper, Junichi YamagishiAbstract:Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a Speech signal without supervision. Until now, the VQ-VAE architecture has previously modeled individual types of Speech features, such as only phones or only F0. This paper introduces an important extension to VQ-VAE for learning F0-related suprasegmental information simultaneously along with traditional phone features.The proposed framework uses two encoders such that the F0 trajectory and Speech Waveform are both input to the system, therefore two separate codebooks are learned. We used a WaveRNN vocoder as the decoder component of VQ-VAE. Our speaker-independent VQ-VAE was trained with raw Speech Waveforms from multi-speaker Japanese Speech databases. Experimental results show that the proposed extension reduces F0 distortion of reconstructed Speech for all unseen test speakers, and results in significantly higher preference scores from a listening test. We additionally conducted experiments using single-speaker Mandarin Speech to demonstrate advantages of our architecture in another language which relies heavily on F0.
-
transferring neural Speech Waveform synthesizers to musical instrument sounds generation
International Conference on Acoustics Speech and Signal Processing, 2020Co-Authors: Yi Zhao, Xin Wang, Lauri Juvela, Junichi YamagishiAbstract:Recent neural Waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in Speech synthesis despite their different methods of Waveform generation. The similarity between Speech and music audio synthesis techniques suggests interesting avenues to explore in terms of the best way to apply Speech synthesizers in the music domain. This work compares three neural synthesizers used for musical instrument sounds generation under three scenarios: training from scratch on music data, zero-shot learning from the Speech domain, and fine-tuning-based adaptation from the Speech to the music domain. The results of a large-scale perceptual test demonstrated that the performance of three synthesizers improved when they were pre-trained on Speech data and fine-tuned on music data, which indicates the usefulness of knowledge from Speech data for music audio generation. Among the synthesizers, WaveGlow showed the best potential in zero-shot learning while NSF performed best in the other scenarios and could generate samples that were perceptually close to natural audio.
-
Training a Neural Speech Waveform Model using Spectral Losses of Short-Time Fourier Transform and Continuous Wavelet Transform
arXiv: Audio and Speech Processing, 2019Co-Authors: Shinji Takaki, Hirokazu Kameoka, Junichi YamagishiAbstract:Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural Speech Waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral amplitude and phase losses obtained by either STFT or continuous wavelet transform (CWT), or both of them. Since CWT is capable of having time and frequency resolutions different from those of STFT and is cable of considering those closer to human auditory scales, the proposed loss functions could provide complementary information on Speech signals. Experimental results showed that it is possible to train a high-quality model by using the proposed CWT spectral loss and is as good as one using STFT-based loss.
-
ICASSP - STFT Spectral Loss for Training a Neural Speech Waveform Model
ICASSP 2019 - 2019 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2019Co-Authors: Shinji Takaki, Toru Nakashika, Xin Wang, Junichi YamagishiAbstract:This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural Speech Waveform model that predicts raw continuous Speech Waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated Speech Waveforms are used to calculate the proposed loss. We also mathematically show that training of the Waveform model on the basis of the proposed loss can be interpreted as maximum likelihood training that assumes the amplitude and phase spectra of generated Speech Waveforms following Gaussian and von Mises distributions, respectively. Furthermore, this paper presents a simple network architecture as the Speech Waveform model, which is composed of uni-directional long short-term memories (LSTMs) and an auto-regressive structure. Experimental results showed that the proposed neural model synthesized high-quality Speech Waveforms.
-
STFT spectral loss for training a neural Speech Waveform model
arXiv: Audio and Speech Processing, 2018Co-Authors: Shinji Takaki, Toru Nakashika, Xin Wang, Junichi YamagishiAbstract:This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural Speech Waveform model that predicts raw continuous Speech Waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated Speech Waveforms are used to calculate the proposed loss. We also mathematically show that training of the Waveform model on the basis of the proposed loss can be interpreted as maximum likelihood training that assumes the amplitude and phase spectra of generated Speech Waveforms following Gaussian and von Mises distributions, respectively. Furthermore, this paper presents a simple network architecture as the Speech Waveform model, which is composed of uni-directional long short-term memories (LSTMs) and an auto-regressive structure. Experimental results showed that the proposed neural model synthesized high-quality Speech Waveforms.
Thomas F. Quatieri - One of the best experts on this subject based on the ideXlab platform.
-
INTERSpeech - Time-varying autoregressive tests for multiscale Speech analysis.
2009Co-Authors: Daniel Rudoy, Thomas F. Quatieri, Patrick J. WolfeAbstract:In this paper we develop hypothesis tests for Speech Waveform nonstationarity based on time-varying autoregressive models, and demonstrate their efficacy in Speech analysis tasks at both segmental and sub-segmental scales. Key to the successful synthesis of these ideas is our employment of a generalized likelihood ratio testing framework tailored to autoregressive coefficient evolutions suitable for Speech. After evaluating our framework on Speech-like synthetic signals, we present preliminary results for two distinct analysis tasks using Speech Waveform data. At the segmental level, we develop an adaptive short-time segmentation scheme and evaluate it on whispered Speech recordings, while at the sub-segmental level, we address the problem of detecting the glottal flow closed phase. Results show that our hypothesis testing framework can reliably detect changes in the vocal tract parameters across multiple scales, thereby underscoring its broad applicability to Speech analysis.
-
Peak-to-RMS reduction of Speech based on a sinusoidal model
IEEE Transactions on Signal Processing, 1991Co-Authors: Thomas F. Quatieri, R.j. McaulayAbstract:A sinusoidal-based analysis/synthesis system is used to apply a radar design solution to the problem of dispersing the phase of a Speech Waveform. Unlike conventional methods of phase dispersion, this solution technique adapts dynamically to the pitch and spectral characteristics of the Speech, while maintaining the original spectral envelope. The solution can also be used to drive the sine-wave amplitude modification for amplitude compression, and is coupled to the desired shaping of the Speech spectrum. The proposed dispersion solution, when integrated with amplitude compression, results in a significant reduction in the peak-to-RMS (root-mean-square) ratio of the Speech Waveform with acceptable loss in quality. Application of a real-time prototype sine-wave preprocessor to AM radio broadcasting is described. >
-
ICASSP - Pitch estimation and voicing detection based on a sinusoidal Speech model
International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: R.j. Mcaulay, Thomas F. QuatieriAbstract:A technique for estimating the pitch of a Speech Waveform is developed. It fits a harmonic set of sine waves to the input data using a mean-squared-error (MSE) criterion. By exploiting a sinusoidal model for the input Speech Waveform, a pitch estimation criterion is derived that is inherently unambiguous, uses pitch-adaptive resolution, uses small-signal suppression to provide enhanced discrimination, and uses amplitude compression to eliminate the effects of pitch-formant interaction. The normalized minimum mean squared error proves to be a powerful discriminant for estimating the likelihood that a given frame of Speech is voiced. >
-
ICASSP - Sinewave-based phase dispersion for audio preprocessing
ICASSP-88. International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: Thomas F. Quatieri, R.j. McaulayAbstract:A sinusoidal-based analysis/synthesis system is used to apply radar signal design solutions to the problem of dispersing the phase of a Speech Waveform. Integrated with dynamic range compression, the resulting system can give a significant reduction in peak/RMS ratio with acceptable loss in quality. The spectral information in the resulting processed Speech Waveform is embedded primarily within the zero crossings of the modified Waveform, rather than the Waveform shape. Consequently, this dispersion technique also serves as a preprocessor for Waveform clipping, allowing considerably deeper thresholding than can be tolerated on the original Waveform. >
Keiichi Tokuda - One of the best experts on this subject based on the ideXlab platform.
-
Mel-Cepstrum-Based Quantization Noise Shaping Applied to Neural-Network-Based Speech Waveform Synthesis
IEEE Transactions on Audio Speech and Language Processing, 2018Co-Authors: Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi TokudaAbstract:This paper presents a mel-cepstrum-based quantization noise shaping method for improving the quality of synthetic Speech generated by neural-network-based Speech Waveform synthesis systems. Since mel-cepstral coefficients closely match the characteristics of human auditory perception, the proposed method effectively masks the white noise introduced by the quantization typically used in neural-network-based Speech Waveform synthesis systems. The paper also describes a computationally efficient implementation of the proposed method using the structure of the mel-log spectrum approximation filter. Experiments using the WaveNet generative model, which is a state-of-the-art model for neural-network-based Speech Waveform synthesis, showed that Speech quality is significantly improved by the proposed method.
-
ICASSP - Directly modeling voiced and unvoiced components in Speech Waveforms by neural networks
2016 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2016Co-Authors: Keiichi Tokuda, Heiga ZenAbstract:This paper proposes a novel acoustic model based on neural networks for statistical parametric Speech synthesis. The neural network outputs parameters of a non-zero mean Gaussian process, which defines a probability density function of a Speech Waveform given linguistic features. The mean and covariance functions of the Gaussian process represent deterministic (voiced) and stochastic (unvoiced) components of a Speech Waveform, whereas the previous approach considered the unvoiced component only. Experimental results show that the proposed approach can generate Speech Waveforms approximating natural Speech Waveforms.
R.j. Mcaulay - One of the best experts on this subject based on the ideXlab platform.
-
Peak-to-RMS reduction of Speech based on a sinusoidal model
IEEE Transactions on Signal Processing, 1991Co-Authors: Thomas F. Quatieri, R.j. McaulayAbstract:A sinusoidal-based analysis/synthesis system is used to apply a radar design solution to the problem of dispersing the phase of a Speech Waveform. Unlike conventional methods of phase dispersion, this solution technique adapts dynamically to the pitch and spectral characteristics of the Speech, while maintaining the original spectral envelope. The solution can also be used to drive the sine-wave amplitude modification for amplitude compression, and is coupled to the desired shaping of the Speech spectrum. The proposed dispersion solution, when integrated with amplitude compression, results in a significant reduction in the peak-to-RMS (root-mean-square) ratio of the Speech Waveform with acceptable loss in quality. Application of a real-time prototype sine-wave preprocessor to AM radio broadcasting is described. >
-
ICASSP - Pitch estimation and voicing detection based on a sinusoidal Speech model
International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: R.j. Mcaulay, Thomas F. QuatieriAbstract:A technique for estimating the pitch of a Speech Waveform is developed. It fits a harmonic set of sine waves to the input data using a mean-squared-error (MSE) criterion. By exploiting a sinusoidal model for the input Speech Waveform, a pitch estimation criterion is derived that is inherently unambiguous, uses pitch-adaptive resolution, uses small-signal suppression to provide enhanced discrimination, and uses amplitude compression to eliminate the effects of pitch-formant interaction. The normalized minimum mean squared error proves to be a powerful discriminant for estimating the likelihood that a given frame of Speech is voiced. >
-
ICASSP - Sinewave-based phase dispersion for audio preprocessing
ICASSP-88. International Conference on Acoustics Speech and Signal Processing, 1Co-Authors: Thomas F. Quatieri, R.j. McaulayAbstract:A sinusoidal-based analysis/synthesis system is used to apply radar signal design solutions to the problem of dispersing the phase of a Speech Waveform. Integrated with dynamic range compression, the resulting system can give a significant reduction in peak/RMS ratio with acceptable loss in quality. The spectral information in the resulting processed Speech Waveform is embedded primarily within the zero crossings of the modified Waveform, rather than the Waveform shape. Consequently, this dispersion technique also serves as a preprocessor for Waveform clipping, allowing considerably deeper thresholding than can be tolerated on the original Waveform. >
Yi Zhao - One of the best experts on this subject based on the ideXlab platform.
-
Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
arXiv: Audio and Speech Processing, 2020Co-Authors: Yi Zhao, Cheng-i Lai, Jennifer Williams, Erica Cooper, Junichi YamagishiAbstract:Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a Speech signal without supervision. Until now, the VQ-VAE architecture has previously modeled individual types of Speech features, such as only phones or only F0. This paper introduces an important extension to VQ-VAE for learning F0-related suprasegmental information simultaneously along with traditional phone features.The proposed framework uses two encoders such that the F0 trajectory and Speech Waveform are both input to the system, therefore two separate codebooks are learned. We used a WaveRNN vocoder as the decoder component of VQ-VAE. Our speaker-independent VQ-VAE was trained with raw Speech Waveforms from multi-speaker Japanese Speech databases. Experimental results show that the proposed extension reduces F0 distortion of reconstructed Speech for all unseen test speakers, and results in significantly higher preference scores from a listening test. We additionally conducted experiments using single-speaker Mandarin Speech to demonstrate advantages of our architecture in another language which relies heavily on F0.
-
transferring neural Speech Waveform synthesizers to musical instrument sounds generation
International Conference on Acoustics Speech and Signal Processing, 2020Co-Authors: Yi Zhao, Xin Wang, Lauri Juvela, Junichi YamagishiAbstract:Recent neural Waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in Speech synthesis despite their different methods of Waveform generation. The similarity between Speech and music audio synthesis techniques suggests interesting avenues to explore in terms of the best way to apply Speech synthesizers in the music domain. This work compares three neural synthesizers used for musical instrument sounds generation under three scenarios: training from scratch on music data, zero-shot learning from the Speech domain, and fine-tuning-based adaptation from the Speech to the music domain. The results of a large-scale perceptual test demonstrated that the performance of three synthesizers improved when they were pre-trained on Speech data and fine-tuned on music data, which indicates the usefulness of knowledge from Speech data for music audio generation. Among the synthesizers, WaveGlow showed the best potential in zero-shot learning while NSF performed best in the other scenarios and could generate samples that were perceptually close to natural audio.
-
Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
2020Co-Authors: Yi Zhao, Li Haoyu, Lai Cheng-i, Williams Jennifer, Cooper Erica, Yamagishi JunichiAbstract:Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a Speech signal without supervision. Until now, the VQ-VAE architecture has previously modeled individual types of Speech features, such as only phones or only F0. This paper introduces an important extension to VQ-VAE for learning F0-related suprasegmental information simultaneously along with traditional phone features.The proposed framework uses two encoders such that the F0 trajectory and Speech Waveform are both input to the system, therefore two separate codebooks are learned. We used a WaveRNN vocoder as the decoder component of VQ-VAE. Our speaker-independent VQ-VAE was trained with raw Speech Waveforms from multi-speaker Japanese Speech databases. Experimental results show that the proposed extension reduces F0 distortion of reconstructed Speech for all unseen test speakers, and results in significantly higher preference scores from a listening test. We additionally conducted experiments using single-speaker Mandarin Speech to demonstrate advantages of our architecture in another language which relies heavily on F0.Comment: Submitted to InterSpeech 202