The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform
Junichi Yamagishi - One of the best experts on this subject based on the ideXlab platform.
-
how similar or different is rakugo Speech Synthesizer to professional performers
International Conference on Acoustics Speech and Signal Processing, 2021Co-Authors: Shuhei Kato, Yusuke Yasuda, Xin Wang, Erica Cooper, Junichi YamagishiAbstract:We have been working on Speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward Speech synthesis that authentically entertains audiences. In this paper, we propose a novel evaluation methodology using synthesized rakugo Speech and real rakugo Speech uttered by professional performers of three different ranks. The naturalness of the synthesized Speech was comparable to that of the human Speech, but the synthesized Speech entertained listeners less than the performers of any rank. However, we obtained some interesting insights into challenges to be solved in order to achieve a truly entertaining rakugo Synthesizer. For example, naturalness was not the most important factor, even though it has generally been emphasized as the most important point to be evaluated in the conventional Speech synthesis field. More important factors were the understandability of the content and distinguishability of the characters in the rakugo story, both of which the synthesized rakugo Speech was relatively inferior at as compared with the professional performers. We also found that fundamental frequency fo modeling should be further improved to better entertain audiences. These results show important steps to reaching authentically entertaining Speech synthesis.
-
detection of synthetic Speech for the problem of imposture
International Conference on Acoustics Speech and Signal Processing, 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to synthetic Speech. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect synthetic Speech and develop a GMM-based synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
-
HMM-based Speech synthesis utilizing glottal inverse filtering
IEEE Transactions on Audio Speech and Language Processing, 2011Co-Authors: Tuomo Raitio, Jussi Nurminen, Antti Suni, Martti Vainio, Junichi Yamagishi, Hannu Pulakka, Paavo AlkuAbstract:This paper describes an hidden Markov model (HMM)-based Speech Synthesizer that utilizes glottal inverse filtering for generating natural sounding synthetic Speech. In the proposed method, Speech is first decomposed into the glottal source signal and the model of the vocal tract filter through glottal inverse filtering, and thus parametrized into excitation and spectral features. The source and filter features are modeled individually in the framework of HMM and generated in the synthesis stage according to the text input. The glottal excitation is synthesized through interpolating and concatenating natural glottal flow pulses, and the excitation signal is further modified according to the spectrum of the desired voice source characteristics. Speech is synthesized by filtering the reconstructed source signal with the vocal tract filter. Experiments show that the proposed system is capable of generating natural sounding Speech, and the quality is clearly better compared to two HMM-based Speech synthesis systems based on widely used vocoder techniques.
-
revisiting the security of speaker verification systems against imposture using synthetic Speech
International Conference on Acoustics Speech and Signal Processing, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Vijendra Raj Apsingekar, Junichi YamagishiAbstract:In this paper, we investigate imposture using synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both speaker verification (SV) and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer which creates synthetic Speech for a targeted speaker through adaptation of a background model. We use two SV systems: standard GMM-UBM-based and a newer SVM-based. Our results show when the systems are tested with human Speech, there are zero false acceptances and zero false rejections. However, when the systems are tested with synthesized Speech, all claims for the targeted speaker are accepted while all other claims are rejected. We propose a two-step process for detection of synthesized Speech in order to prevent this imposture. Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech will lead to an unacceptably high false acceptance rate.
-
evaluation of the vulnerability of speaker verification to synthetic Speech
Odyssey, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Junichi YamagishiAbstract:In this paper, we evaluate the vulnerability of a speaker verification (SV) system to synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both SV and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and a GMM-UBM-based SV system. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV system has a 0.4% EER. When the system is tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, 90% of the matched claims are accepted. This result suggests a possible vulnerability in SV systems to synthetic Speech. In order to detect synthetic Speech prior to recognition, we investigate the use of an automatic Speech recognizer (ASR), dynamic-timewarping (DTW) distance of mel-frequency cepstral coefficients (MFCC), and previously-proposed average inter-frame difference of log-likelihood (IFDLL). Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech can lead to an unacceptably high acceptance rate of synthetic speakers.
Phillip L. De Leon - One of the best experts on this subject based on the ideXlab platform.
-
detection of synthetic Speech for the problem of imposture
International Conference on Acoustics Speech and Signal Processing, 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to synthetic Speech. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect synthetic Speech and develop a GMM-based synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
-
revisiting the security of speaker verification systems against imposture using synthetic Speech
International Conference on Acoustics Speech and Signal Processing, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Vijendra Raj Apsingekar, Junichi YamagishiAbstract:In this paper, we investigate imposture using synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both speaker verification (SV) and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer which creates synthetic Speech for a targeted speaker through adaptation of a background model. We use two SV systems: standard GMM-UBM-based and a newer SVM-based. Our results show when the systems are tested with human Speech, there are zero false acceptances and zero false rejections. However, when the systems are tested with synthesized Speech, all claims for the targeted speaker are accepted while all other claims are rejected. We propose a two-step process for detection of synthesized Speech in order to prevent this imposture. Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech will lead to an unacceptably high false acceptance rate.
-
evaluation of the vulnerability of speaker verification to synthetic Speech
Odyssey, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Junichi YamagishiAbstract:In this paper, we evaluate the vulnerability of a speaker verification (SV) system to synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both SV and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and a GMM-UBM-based SV system. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV system has a 0.4% EER. When the system is tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, 90% of the matched claims are accepted. This result suggests a possible vulnerability in SV systems to synthetic Speech. In order to detect synthetic Speech prior to recognition, we investigate the use of an automatic Speech recognizer (ASR), dynamic-timewarping (DTW) distance of mel-frequency cepstral coefficients (MFCC), and previously-proposed average inter-frame difference of log-likelihood (IFDLL). Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech can lead to an unacceptably high acceptance rate of synthetic speakers.
Thierry Dutoit - One of the best experts on this subject based on the ideXlab platform.
-
a deterministic plus stochastic model of the residual signal for improved parametric Speech synthesis
arXiv: Sound, 2019Co-Authors: Thomas Drugman, Geoffrey Wilfart, Thierry DutoitAbstract:Speech generated by parametric Synthesizers generally suffers from a typical buzziness, similar to what was encountered in old LPC-like vocoders. In order to alleviate this problem, a more suited modeling of the excitation should be adopted. For this, we hereby propose an adaptation of the Deterministic plus Stochastic Model (DSM) for the residual. In this model, the excitation is divided into two distinct spectral bands delimited by the maximum voiced frequency. The deterministic part concerns the low-frequency contents and consists of a decomposition of pitch-synchronous residual frames on an orthonormal basis obtained by Principal Component Analysis. The stochastic component is a high-pass filtered noise whose time structure is modulated by an energy-envelope, similarly to what is done in the Harmonic plus Noise Model (HNM). The proposed residual model is integrated within a HMM-based Speech Synthesizer and is compared to the traditional excitation through a subjective test. Results show a significative improvement for both male and female voices. In addition the proposed model requires few computational load and memory, which is essential for its integration in commercial applications.
-
a deterministic plus stochastic model of the residual signal for improved parametric Speech synthesis
Conference of the International Speech Communication Association, 2009Co-Authors: Thomas Drugman, Geoffrey Wilfart, Thierry DutoitAbstract:Speech generated by parametric Synthesizers generally suffers from a typical buzziness, similar to what was encountered in old LPC-like vocoders. In order to alleviate this problem, a more suited modeling of the excitation should be adopted. For this, we hereby propose an adaptation of the Deterministic plus Stochastic Model (DSM) for the residual. In this model, the excitation is divided into two distinct spectral bands delimited by the maximum voiced frequency. The deterministic part concerns the low-frequency contents and consists of a decomposition of pitch-synchronous residual frames on an orthonormal basis obtained by Principal Component Analysis. The stochastic component is a high-pass filtered noise whose time structure is modulated by an energy-envelope, similarly to what is done in the Harmonic plus Noise Model (HNM). The proposed residual model is integrated within a HMM-based Speech Synthesizer and is compared to the traditional excitation through a subjective test. Results show a significative improvement for both male and female voices. In addition the proposed model requires few computational load and memory, which is essential for its integration in commercial applications. Index Terms: HMM-based Speech synthesis, residual modeling, Deterministic plus Stochastic model
-
high quality Speech synthesis for phonetic Speech segmentation
Conference of the International Speech Communication Association, 1997Co-Authors: Fabrice Malfrere, Thierry DutoitAbstract:This paper presents an original technique for solving the phonetic segmentation problem. It is based on the use of a Speech Synthesizer for the alignment of a text on its corresponding Speech signal. A high-quality digital Speech Synthesizer is used to create a synthetic reference Speech pattern used in the alignment process. This approach has the great advantage on other approaches that no training stage (hence no labeled database) is needed. The system has been mainly evaluated on French read utterances. Other evaluations have been made on other languages like English, German, Romanian and Spanish. Following these experiments, the system seems to be a powerful tool for the automatic constitution of large phonetically and prosodically labeled Speech databases. The availability of such corpora will be a key point for the development of improved Speech synthesis and recognition systems.
-
the mbrola project towards a set of high quality Speech Synthesizers free of use for non commercial purposes
International Conference on Spoken Language Processing, 1996Co-Authors: Thierry Dutoit, Vincent Pagel, Nicolas Pierret, F Bataille, O Van Der VreckenAbstract:The aim of the MBROLA project, initiated by the Faculte Polytechnique de Mons (Belgium), is to obtain a set of Speech Synthesizers for as many voices, languages and dialects as possible, free of use for non-commercial and non-military applications. The ultimate goal is to boost academic research on Speech synthesis, and particularly on prosody generation, known as one of the biggest challenges taken up by text-to-Speech Synthesizers for the years to come. Central to the MBROLA project is MBROLA 2.00, a Speech Synthesizer based on the concatenation of diphones. Executable files of this Synthesizer have been made freely available for many computers/operating systems, as well as a first diphone database for a French male voice. We describe the terms of participation to the project, as a user, as an associated developer, or as a database provider.
-
mbr psola text to Speech synthesis based on an mbe re synthesis of the segments database
Speech Communication, 1993Co-Authors: Thierry Dutoit, H LeichAbstract:Abstract The use of the Time-Domain Pitch Synchronous OverLap-Add (TD-PSOLA) algorithm in a Text-To-Speech Synthesizer is reviewed. Its drawbacks are underlined and three conditions on the Speech database are examined. In order to satisfy them, a previously described high quality resynthesis process is developed and enhanced, which makes use of the well-known Multi-Band Excited (MBE) model. An important by-product of this operation is that optimal Pitch Marking turns out to be automatic. A temporal interpolation block is finally added. The resulting Multi-Band Resynthesis Pitch Synchronous OverLap Add (MBR-PSOLA) synthesis algorithm supports spectral interpolation between voiced parts of segments, with virtually no increase in complexity. It provides the basis of a high-quality Text-To-Speech (TTS) Synthesizer.
Michael Pucher - One of the best experts on this subject based on the ideXlab platform.
-
detection of synthetic Speech for the problem of imposture
International Conference on Acoustics Speech and Signal Processing, 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to synthetic Speech. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect synthetic Speech and develop a GMM-based synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
-
revisiting the security of speaker verification systems against imposture using synthetic Speech
International Conference on Acoustics Speech and Signal Processing, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Vijendra Raj Apsingekar, Junichi YamagishiAbstract:In this paper, we investigate imposture using synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both speaker verification (SV) and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer which creates synthetic Speech for a targeted speaker through adaptation of a background model. We use two SV systems: standard GMM-UBM-based and a newer SVM-based. Our results show when the systems are tested with human Speech, there are zero false acceptances and zero false rejections. However, when the systems are tested with synthesized Speech, all claims for the targeted speaker are accepted while all other claims are rejected. We propose a two-step process for detection of synthesized Speech in order to prevent this imposture. Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech will lead to an unacceptably high false acceptance rate.
-
evaluation of the vulnerability of speaker verification to synthetic Speech
Odyssey, 2010Co-Authors: Phillip L. De Leon, Michael Pucher, Junichi YamagishiAbstract:In this paper, we evaluate the vulnerability of a speaker verification (SV) system to synthetic Speech. Although this problem was first examined over a decade ago, dramatic improvements in both SV and Speech synthesis have renewed interest in this problem. We use a HMM-based Speech Synthesizer, which creates synthetic Speech for a targeted speaker through adaptation of a background model and a GMM-UBM-based SV system. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV system has a 0.4% EER. When the system is tested with synthetic Speech generated from speaker models derived from the WSJ journal corpus, 90% of the matched claims are accepted. This result suggests a possible vulnerability in SV systems to synthetic Speech. In order to detect synthetic Speech prior to recognition, we investigate the use of an automatic Speech recognizer (ASR), dynamic-timewarping (DTW) distance of mel-frequency cepstral coefficients (MFCC), and previously-proposed average inter-frame difference of log-likelihood (IFDLL). Overall, while SV systems have impressive accuracy, even with the proposed detector, high-quality synthetic Speech can lead to an unacceptably high acceptance rate of synthetic speakers.
Frank H Guenther - One of the best experts on this subject based on the ideXlab platform.
-
Role of the auditory system in Speech production.
Handbook of clinical neurology, 2015Co-Authors: Frank H Guenther, Gregory HickokAbstract:This chapter reviews evidence regarding the role of auditory perception in shaping Speech output. Evidence indicates that Speech movements are planned to follow auditory trajectories. This in turn is followed by a description of the Directions Into Velocities of Articulators (DIVA) model, which provides a detailed account of the role of auditory feedback in Speech motor development and control. A brief description of the higher-order brain areas involved in Speech sequencing (including the pre-supplementary motor area and inferior frontal sulcus) is then provided, followed by a description of the Hierarchical State Feedback Control (HSFC) model, which posits internal error detection and correction processes that can detect and correct Speech production errors prior to articulation. The chapter closes with a treatment of promising future directions of research into auditory-motor interactions in Speech, including the use of intracranial recording techniques such as electrocorticography in humans, the investigation of the potential roles of various large-scale brain rhythms in Speech perception and production, and the development of brain-computer interfaces that use auditory feedback to allow profoundly paralyzed users to learn to produce Speech using a Speech Synthesizer.
-
Artificial Speech Synthesizer control by brain-computer interface
Proceedings of the Annual Conference of the International Speech Communication Association INTERSPEECH, 2009Co-Authors: Jonathan S. Brumberg, Philip R Kennedy, Frank H GuentherAbstract:We developed and tested a brain-computer interface for control of an artificial Speech Synthesizer by an individual with near complete paralysis. This neural prosthesis for Speech restoration is currently capable of predicting vowel formant frequencies based on neural activity recorded from an intracortical microelectrode implanted in the left hemisphere Speech motor cortex. Using instantaneous auditory feedback (< 50 ms) of predicted formant frequencies, the study participant has been able to correctly perform a vowel production task at a maximum rate of 80-90 % correct. Index Terms: Speech synthesis, brain computer interface 1.