The Experts below are selected from a list of 22215 Experts worldwide ranked by ideXlab platform
Tomoki Toda - One of the best experts on this subject based on the ideXlab platform.
-
quasi periodic wavenet an autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
IEEE Transactions on Audio Speech and Language Processing, 2021Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the limited pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution neural networks (PDCNNs). Specifically, as a probabilistic autoregressive Generation model with stacked dilated convolution layers, WN achieves high-fidelity audio waveform Generation. However, the pure-data-driven nature and the lack of prior knowledge of audio signals degrade the pitch controllability of WN. For instance, it is difficult for WN to precisely generate the periodic components of audio signals when the given auxiliary fundamental frequency ( $F_{0}$ ) features are outside the $F_{0}$ range observed in the training data. To address this problem, QPNet with two novel designs is proposed. First, the PDCNN component is applied to dynamically change the network architecture of WN according to the given auxiliary $F_{0}$ features. Second, a cascaded network structure is utilized to simultaneously model the long- and short-term dependencies of quasi-periodic signals such as Speech. The performances of single-tone sinusoid and Speech Generations are evaluated. The experimental results show the effectiveness of the PDCNNs for unseen auxiliary $F_{0}$ features and the effectiveness of the cascaded structure for Speech Generation.
-
quasi periodic parallel wavegan a non autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
IEEE Transactions on Audio Speech and Language Processing, 2021Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ( $F_{0}$ ) feature such as a scaled $F_{0}$ . To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.
-
quasi periodic parallel wavegan vocoder a non autoregressive pitch dependent dilated convolution model for parametric Speech Generation
Conference of the International Speech Communication Association, 2020Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) Speech Generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate Speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity Speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated Speech degrades because of the fixed and generic network of PWG without prior knowledge of Speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and Speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable Speech quality of QPPWG-generated Speech when the QPPWG model size is only 70 % of that of vanilla PWG.
-
quasi periodic wavenet an autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
arXiv: Audio and Speech Processing, 2020Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution neural networks (PDCNNs). Specifically, as a probabilistic autoregressive Generation model with stacked dilated convolution layers, WN achieves high-fidelity audio waveform Generation. However, the pure-data-driven nature and the lack of prior knowledge of audio signals degrade the pitch controllability of WN. For instance, it is difficult for WN to precisely generate the periodic components of audio signals when the given auxiliary fundamental frequency (F0) features are outside the F0 range observed in the training data. To address this problem, QPNet with two novel designs is proposed. First, the PDCNN component is applied to dynamically change the network architecture of WN according to the given auxiliary F0 features. Second, a cascaded network structure is utilized to simultaneously model the long- and short-term dependences of quasi-periodic signals such as Speech. The performances of single-tone sinusoid and Speech Generations are evaluated. The experimental results show the effectiveness of the PDCNNs for unseen auxiliary F0 features and the effectiveness of the cascaded structure for Speech Generation.
-
quasi periodic wavenet vocoder a pitch dependent dilated convolution model for parametric Speech Generation
Conference of the International Speech Communication Association, 2019Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, we propose a quasi-periodic neural network (QPNet) vocoder with a novel network architecture named pitch-dependent dilated convolution (PDCNN) to improve the pitch controllability of WaveNet (WN) vocoder. The effectiveness of the WN vocoder to generate high-fidelity Speech samples from given acoustic features has been proved recently. However, because of the fixed dilated convolution and generic network architecture, the WN vocoder hardly generates Speech with given F0 values which are outside the range observed in training data. Consequently, the WN vocoder lacks the pitch controllability which is one of the essential capabilities of conventional vocoders. To address this limitation, we propose the PDCNN component which has the time-variant adaptive dilation size related to the given F0 values and a cascade network structure of the QPNet vocoder to generate quasi-periodic signals such as Speech. Both objective and subjective tests are conducted, and the experimental results demonstrate the better pitch controllability of the QPNet vocoder compared to the same and double sized WN vocoders while attaining comparable Speech qualities. Index Terms: WaveNet, vocoder, quasi-periodic signal, pitch-dependent dilated convolution, pitch controllability
Tomoki Hayashi - One of the best experts on this subject based on the ideXlab platform.
-
quasi periodic wavenet an autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
IEEE Transactions on Audio Speech and Language Processing, 2021Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the limited pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution neural networks (PDCNNs). Specifically, as a probabilistic autoregressive Generation model with stacked dilated convolution layers, WN achieves high-fidelity audio waveform Generation. However, the pure-data-driven nature and the lack of prior knowledge of audio signals degrade the pitch controllability of WN. For instance, it is difficult for WN to precisely generate the periodic components of audio signals when the given auxiliary fundamental frequency ( $F_{0}$ ) features are outside the $F_{0}$ range observed in the training data. To address this problem, QPNet with two novel designs is proposed. First, the PDCNN component is applied to dynamically change the network architecture of WN according to the given auxiliary $F_{0}$ features. Second, a cascaded network structure is utilized to simultaneously model the long- and short-term dependencies of quasi-periodic signals such as Speech. The performances of single-tone sinusoid and Speech Generations are evaluated. The experimental results show the effectiveness of the PDCNNs for unseen auxiliary $F_{0}$ features and the effectiveness of the cascaded structure for Speech Generation.
-
quasi periodic parallel wavegan a non autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
IEEE Transactions on Audio Speech and Language Processing, 2021Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ( $F_{0}$ ) feature such as a scaled $F_{0}$ . To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.
-
quasi periodic parallel wavegan vocoder a non autoregressive pitch dependent dilated convolution model for parametric Speech Generation
Conference of the International Speech Communication Association, 2020Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) Speech Generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate Speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity Speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated Speech degrades because of the fixed and generic network of PWG without prior knowledge of Speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and Speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable Speech quality of QPPWG-generated Speech when the QPPWG model size is only 70 % of that of vanilla PWG.
-
quasi periodic wavenet an autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
arXiv: Audio and Speech Processing, 2020Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution neural networks (PDCNNs). Specifically, as a probabilistic autoregressive Generation model with stacked dilated convolution layers, WN achieves high-fidelity audio waveform Generation. However, the pure-data-driven nature and the lack of prior knowledge of audio signals degrade the pitch controllability of WN. For instance, it is difficult for WN to precisely generate the periodic components of audio signals when the given auxiliary fundamental frequency (F0) features are outside the F0 range observed in the training data. To address this problem, QPNet with two novel designs is proposed. First, the PDCNN component is applied to dynamically change the network architecture of WN according to the given auxiliary F0 features. Second, a cascaded network structure is utilized to simultaneously model the long- and short-term dependences of quasi-periodic signals such as Speech. The performances of single-tone sinusoid and Speech Generations are evaluated. The experimental results show the effectiveness of the PDCNNs for unseen auxiliary F0 features and the effectiveness of the cascaded structure for Speech Generation.
-
quasi periodic wavenet vocoder a pitch dependent dilated convolution model for parametric Speech Generation
Conference of the International Speech Communication Association, 2019Co-Authors: Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki TodaAbstract:In this paper, we propose a quasi-periodic neural network (QPNet) vocoder with a novel network architecture named pitch-dependent dilated convolution (PDCNN) to improve the pitch controllability of WaveNet (WN) vocoder. The effectiveness of the WN vocoder to generate high-fidelity Speech samples from given acoustic features has been proved recently. However, because of the fixed dilated convolution and generic network architecture, the WN vocoder hardly generates Speech with given F0 values which are outside the range observed in training data. Consequently, the WN vocoder lacks the pitch controllability which is one of the essential capabilities of conventional vocoders. To address this limitation, we propose the PDCNN component which has the time-variant adaptive dilation size related to the given F0 values and a cascade network structure of the QPNet vocoder to generate quasi-periodic signals such as Speech. Both objective and subjective tests are conducted, and the experimental results demonstrate the better pitch controllability of the QPNet vocoder compared to the same and double sized WN vocoders while attaining comparable Speech qualities. Index Terms: WaveNet, vocoder, quasi-periodic signal, pitch-dependent dilated convolution, pitch controllability
Hisashi Kawai - One of the best experts on this subject based on the ideXlab platform.
-
quasi periodic parallel wavegan a non autoregressive raw waveform generative model with pitch dependent dilated convolution neural network
IEEE Transactions on Audio Speech and Language Processing, 2021Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ( $F_{0}$ ) feature such as a scaled $F_{0}$ . To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.
-
quasi periodic parallel wavegan vocoder a non autoregressive pitch dependent dilated convolution model for parametric Speech Generation
Conference of the International Speech Communication Association, 2020Co-Authors: Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki TodaAbstract:In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) Speech Generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate Speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity Speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated Speech degrades because of the fixed and generic network of PWG without prior knowledge of Speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and Speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable Speech quality of QPPWG-generated Speech when the QPPWG model size is only 70 % of that of vanilla PWG.
-
an investigation of noise shaping with perceptual weighting for wavenet based Speech Generation
International Conference on Acoustics Speech and Signal Processing, 2018Co-Authors: Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi KawaiAbstract:We propose a noise shaping method to improve the sound quality of Speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by the quantization error generated by representing waveform samples as discrete symbols and the prediction error of the CNN. We analyze these noise signals and show that 1) since the prediction error is much larger than the quantization error, the effect of the quantization error on the noise signals is practically negligible, and 2) noise signals tend to cause large spectral distortion in a high-frequency band. To alleviate the adverse effect of these noise signals on the generated Speech signals, the proposed noise shaping method applies a perceptual weighting filter to WaveNet, making it possible to use the frequency masking properties of the human auditory system. We conducted objective and subjective evaluations to investigate the effectiveness of the proposed method and demonstrated that it significantly improved the sound quality of the generated Speech signals.
Toda Tomoki - One of the best experts on this subject based on the ideXlab platform.
-
Quasi-Periodic Parallel WaveGAN: A Non-autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
'Institute of Electrical and Electronics Engineers (IEEE)', 2021Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ($F_{0}$) feature such as a scaled $F_{0}$. To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.Comment: 15 pages, 10 figures, 8 table
-
Quasi-Periodic Parallel WaveGAN Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model for Parametric Speech Generation
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) Speech Generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate Speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity Speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated Speech degrades because of the fixed and generic network of PWG without prior knowledge of Speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and Speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable Speech quality of QPPWG-generated Speech when the QPPWG model size is only 70 % of that of vanilla PWG.Comment: 5 page, 6 figures, 2 tables. Proc. InterSpeech, 202
-
Quasi-Periodic Parallel WaveGAN: A Non-autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ($F_{0}$) feature such as a scaled $F_{0}$. To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.Comment: 13 pages, 10 figures, 8 tables. Submitted to a journa
-
Quasi-Periodic WaveNet Vocoder: A Pitch Dependent Dilated Convolution Model for Parametric Speech Generation
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Tobing, Patrick Lumban, Kobayashi Kazuhiro, Toda TomokiAbstract:In this paper, we propose a quasi-periodic neural network (QPNet) vocoder with a novel network architecture named pitch-dependent dilated convolution (PDCNN) to improve the pitch controllability of WaveNet (WN) vocoder. The effectiveness of the WN vocoder to generate high-fidelity Speech samples from given acoustic features has been proved recently. However, because of the fixed dilated convolution and generic network architecture, the WN vocoder hardly generates Speech with given F0 values which are outside the range observed in training data. Consequently, the WN vocoder lacks the pitch controllability which is one of the essential capabilities of conventional vocoders. To address this limitation, we propose the PDCNN component which has the time-variant adaptive dilation size related to the given F0 values and a cascade network structure of the QPNet vocoder to generate quasi-periodic signals such as Speech. Both objective and subjective tests are conducted, and the experimental results demonstrate the better pitch controllability of the QPNet vocoder compared to the same and double sized WN vocoders while attaining comparable Speech qualities. Index Terms: WaveNet, vocoder, quasi-periodic signal, pitch-dependent dilated convolution, pitch controllabilityComment: 5 pages, 4 figures, Proc. InterSpeech, 201
-
Collapsed Speech segment detection and suppression for WaveNet vocoder
2018Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Tobing, Patrick Lumban, Kobayashi Kazuhiro, Toda TomokiAbstract:In this paper, we propose a technique to alleviate the quality degradation caused by collapsed Speech segments sometimes generated by the WaveNet vocoder. The effectiveness of the WaveNet vocoder for generating natural Speech from acoustic features has been proved in recent works. However, it sometimes generates very noisy Speech with collapsed Speech segments when only a limited amount of training data is available or significant acoustic mismatches exist between the training and testing data. Such a limitation on the corpus and limited ability of the model can easily occur in some Speech Generation applications, such as voice conversion and Speech enhancement. To address this problem, we propose a technique to automatically detect collapsed Speech segments. Moreover, to refine the detected segments, we also propose a waveform Generation technique for WaveNet using a linear predictive coding constraint. Verification and subjective tests are conducted to investigate the effectiveness of the proposed techniques. The verification results indicate that the detection technique can detect most collapsed segments. The subjective evaluations of voice conversion demonstrate that the Generation technique significantly improves the Speech quality while maintaining the same speaker similarity.Comment: 5 pages, 6 figures. Proc. InterSpeech, 201
Wu Yi-chiao - One of the best experts on this subject based on the ideXlab platform.
-
Quasi-Periodic Parallel WaveGAN: A Non-autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
'Institute of Electrical and Electronics Engineers (IEEE)', 2021Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ($F_{0}$) feature such as a scaled $F_{0}$. To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.Comment: 15 pages, 10 figures, 8 table
-
Quasi-Periodic Parallel WaveGAN Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model for Parametric Speech Generation
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) Speech Generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate Speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity Speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated Speech degrades because of the fixed and generic network of PWG without prior knowledge of Speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and Speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable Speech quality of QPPWG-generated Speech when the QPPWG model size is only 70 % of that of vanilla PWG.Comment: 5 page, 6 figures, 2 tables. Proc. InterSpeech, 202
-
Quasi-Periodic Parallel WaveGAN: A Non-autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Okamoto Takuma, Kawai Hisashi, Toda TomokiAbstract:In this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose Generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity Speech Generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency ($F_{0}$) feature such as a scaled $F_{0}$. To improve the pitch controllability and Speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary $F_{0}$ feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary $F_{0}$ feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure.Comment: 13 pages, 10 figures, 8 tables. Submitted to a journa
-
Quasi-Periodic WaveNet Vocoder: A Pitch Dependent Dilated Convolution Model for Parametric Speech Generation
2020Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Tobing, Patrick Lumban, Kobayashi Kazuhiro, Toda TomokiAbstract:In this paper, we propose a quasi-periodic neural network (QPNet) vocoder with a novel network architecture named pitch-dependent dilated convolution (PDCNN) to improve the pitch controllability of WaveNet (WN) vocoder. The effectiveness of the WN vocoder to generate high-fidelity Speech samples from given acoustic features has been proved recently. However, because of the fixed dilated convolution and generic network architecture, the WN vocoder hardly generates Speech with given F0 values which are outside the range observed in training data. Consequently, the WN vocoder lacks the pitch controllability which is one of the essential capabilities of conventional vocoders. To address this limitation, we propose the PDCNN component which has the time-variant adaptive dilation size related to the given F0 values and a cascade network structure of the QPNet vocoder to generate quasi-periodic signals such as Speech. Both objective and subjective tests are conducted, and the experimental results demonstrate the better pitch controllability of the QPNet vocoder compared to the same and double sized WN vocoders while attaining comparable Speech qualities. Index Terms: WaveNet, vocoder, quasi-periodic signal, pitch-dependent dilated convolution, pitch controllabilityComment: 5 pages, 4 figures, Proc. InterSpeech, 201
-
Collapsed Speech segment detection and suppression for WaveNet vocoder
2018Co-Authors: Wu Yi-chiao, Hayashi Tomoki, Tobing, Patrick Lumban, Kobayashi Kazuhiro, Toda TomokiAbstract:In this paper, we propose a technique to alleviate the quality degradation caused by collapsed Speech segments sometimes generated by the WaveNet vocoder. The effectiveness of the WaveNet vocoder for generating natural Speech from acoustic features has been proved in recent works. However, it sometimes generates very noisy Speech with collapsed Speech segments when only a limited amount of training data is available or significant acoustic mismatches exist between the training and testing data. Such a limitation on the corpus and limited ability of the model can easily occur in some Speech Generation applications, such as voice conversion and Speech enhancement. To address this problem, we propose a technique to automatically detect collapsed Speech segments. Moreover, to refine the detected segments, we also propose a waveform Generation technique for WaveNet using a linear predictive coding constraint. Verification and subjective tests are conducted to investigate the effectiveness of the proposed techniques. The verification results indicate that the detection technique can detect most collapsed segments. The subjective evaluations of voice conversion demonstrate that the Generation technique significantly improves the Speech quality while maintaining the same speaker similarity.Comment: 5 pages, 6 figures. Proc. InterSpeech, 201