The Experts below are selected from a list of 258 Experts worldwide ranked by ideXlab platform
David B. Pisoni - One of the best experts on this subject based on the ideXlab platform.
-
Perception and Comprehension of Synthetic Speech 1
2004Co-Authors: Stephen J. Winters, David B. PisoniAbstract:An extensive body of research on the perception of Synthetic Speech carried out over the past 30 years has established that listeners have much more difficulty perceiving Synthetic Speech than natural Speech. Differences in perceptual processing have been found in a variety of behavioral tasks, including assessments of segmental intelligibility, word recall, lexical decision, sentence transcription, and comprehension of spoken passages of connected text. Alternative groups of listeners—such as non-native speakers of English, children and older adults—have even more difficulty perceiving Synthetic Speech than young, healthy, college-aged listeners typically tested in perception studies. It has also been shown, however, that the ability to perceive Synthetic Speech improves rapidly with training and experience. Incorporating appropriate prosodic contours into Synthetic Speech algorithms—along with providing listeners with higherlevel contextual information—can also aid the perception of Synthetic Speech. Listener difficulty in processing Synthetic Speech has been attributed to the impoverished acousticphonetic segmental cues—and inherent lack of natural variability and acoustic-phonetic redundancy—in Synthetic Speech produced by rule. The perceptual difficulties that listeners have in perceiving Speech which lacks acoustic-phonetic variability has been cited as evidence for the importance of variability to the perception of natural Speech. Future research on the perception of Synthetic Speech will need to investigate the sources of acoustic-phonetic variability and redundancy that improve the perception of Synthetic Speech, as well as determine the efficacy of Synthetically produced audio-visual Speech, and the extent to which the impoverished acoustic-phonetic structure of Synthetic Speech impacts higher-level comprehension processes. New behavioral methods of assessing the perception of Speech by human listeners will need to be developed in order for our understanding of Synthetic Speech perception to keep pace with the rapid progress of Speech synthesis technology.
-
Perception of Synthetic Speech
Progress in Speech Synthesis, 1997Co-Authors: David B. PisoniAbstract:This chapter summarizes the results we obtained over the last 15 years at Indiana University on the perception of Synthetic Speech produced by rule. A wide variety of behavioral studies have been carried out on phoneme intelligibility, word recognition, and comprehension to learn more about how human listeners perceive and understand Synthetic Speech. Some of this research, particularly the earlier studies on segmental intelligibility, was directed toward applied issues dealing with perceptual evaluation and assessment of different synthesis systems. Other aspects of the research program have been more theoretically motivated and were designed to learn more about Speech perception and spoken language comprehension. Our findings have shown that the perception of Synthetic Speech depends on several general factors including the acoustic-phonetic properties of the Speech signal, the specific cognitive demands of the information-processing task the listener is asked to perform, and the previous background and experience of the listener. Suggestions for future research on improving naturalness, intelligibility, and comprehension are offered in light of several recent findings on the role of stimulus variability and the contribution of indexical factors to Speech perception and spoken word recognition. Our perceptual findings have shown the importance of behavioral testing with human listeners as an integral component of evaluation and assessment techniques in synthesis research and development.
-
Comprehension of Synthetic Speech produced by rule: a review and theoretical interpretation.
Language and Speech, 1992Co-Authors: Susan A. Duffy, David B. PisoniAbstract:In this paper, we review research on the perception and comprehension of Synthetic Speech produced by rule. We discuss the difficulties that Synthetic Speech causes for the listener and the evidence that the immediate result of those difficulties is a delay in the point at which words are recognized. We then argue that this delay in processing affects not only lexical access but also comprehension processes. We consider the mechanisms by which the comprehension system adjusts to this delay, the resulting costs to higher level comprehension processes, and the changes that occur in the language processing system as its familiarity with Synthetic Speech increases. Based on the framework we have developed, we suggest several directions for future research on the comprehension of Synthetic Speech.
-
Comprehension of Synthetic Speech produced by rule: word monitoring and sentence-by-sentence listening times
Human factors, 1991Co-Authors: James V. Ralston, David B. Pisoni, Beth G. Greene, Scott E. Lively, John W. MullennixAbstract:Previous comprehension studies using postperceptual memory tests have often reported negligible differences in performance between natural Speech and several kinds of Synthetic Speech produced by rule, despite large differences in segmental intelligibility. The present experiments investigated the comprehension of natural and Synthetic Speech using two different on-line tasks: word monitoring and sentence-by-sentence listening. On-line task performance was slower and less accurate for passages of Synthetic Speech than for passages of natural Speech. Recognition memory performance in both experiments was less accurate following passages of Synthetic Speech than of natural Speech. Monitoring performance, sentence listening times, and recognition memory accuracy all showed moderate correlations with intelligibility scores obtained using the Modified Rhyme Test. The results suggest that poorer comprehension of passages of Synthetic Speech is attributable in part to the greater encoding demands of Synthetic Speech. In contrast to earlier studies, the present results demonstrate that on-line tasks can be used to measure differences in comprehension performance between natural and Synthetic Speech. Language: en
-
Recognition of Synthetic Speech by hearing-impaired elderly listeners.
Journal of speech and hearing research, 1991Co-Authors: Larry E. Humes, Kathleen J. Nelson, David B. PisoniAbstract:The Modified Rhyme Test (MRT), recorded using natural Speech and two forms of Synthetic Speech, DECtalk and Votrax, was used to measure both open-set and closed-set Speech-recognition performance. ...
Phillip L. De Leon - One of the best experts on this subject based on the ideXlab platform.
-
ICASSP - Performance of I-vector speaker verification and the detection of Synthetic Speech
2014 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2014Co-Authors: Richard D. Mcclanahan, Bryan Stewart, Phillip L. De LeonAbstract:In this paper, we present new research results on the vulnerability of speaker verification (SV) systems to Synthetic Speech. Using a state-of-the-art i-vector SV system and evaluating with the Wall-Street Journal (WSJ) corpus, our SV system has a 0.00% false rejection rate (FRR) and 1.74 × 10 -5 false acceptance rate (FAR). When the i-vector system is tested with state-of-the-art speaker-adaptive, hidden Markov model (HMM)-based Synthetic Speech generated from speaker models derived from the WSJ journal corpus, 22.9% of the matched claims are accepted highlighting the vulnerability of SV systems to Synthetic Speech. We propose a new Synthetic Speech detector (SSD) which uses previously-proposed features derived from image analysis of pitch patterns but extracted on phoneme-level segments and which leverages the available enrollment Speech from the SV system. When the SSD is applied to human and Synthetic Speech accepted by the SV system, the overall system has a FRR of 7.35% and a FAR of 2.34 × 10 -4 which is lower than previously-reported systems and thus significantly reduces the vulnerability.
-
ICASSP - Synthetic Speech detection based on selectedword discriminators
2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013Co-Authors: Phillip L. De Leon, Bryan StewartAbstract:Speaker verification (SV) systems have been shown to be vulnerable to imposture using Speech synthesizers. In this paper, we extend previous work in detecting Synthetic Speech by analyzing words which provide strong discrimination between human and Synthetic Speech. The research is applicable to authentication systems based on text-dependent SV where the user is prompted to speak a certain utterance which can be chosen by the designer. Our results show that this approach to Synthetic Speech detection leads to higher accuracies than other proposed approaches. Using various corpora to train and test, our results show 98% accuracy in correctly classifying both human and Synthetic Speech.
-
evaluation of speaker verification security and detection of hmm based Synthetic Speech
IEEE Transactions on Audio Speech and Language Processing, 2012Co-Authors: Phillip L. De Leon, Inma Hernaez, Michael Pucher, Junichi Yamagishi, Ibon SaratxagaAbstract:In this paper, we evaluate the vulnerability of speaker verification (SV) systems to Synthetic Speech. The SV systems are based on either the Gaussian mixture model–universal background model (GMM-UBM) or support vector machine (SVM) using GMM supervectors. We use a hidden Markov model (HMM)-based text-to-Speech (TTS) synthesizer, which can synthesize Speech for a target speaker using small amounts of training data through model adaptation of an average voice or background model. Although the SV systems have a very low equal error rate (EER), when tested with Synthetic Speech generated from speaker models derived from the Wall Street Journal (WSJ) Speech corpus, over 81% of the matched claims are accepted. This result suggests vulnerability in SV systems and thus a need to accurately detect Synthetic Speech. We propose a new feature based on relative phase shift (RPS), demonstrate reliable detection of Synthetic Speech, and show how this classifier can be used to improve security of SV systems.
-
Synthetic Speech discrimination using pitch pattern statistics derived from image analysis
Conference of the International Speech Communication Association, 2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
-
INTERSpeech - Synthetic Speech Discrimination using Pitch Pattern Statistics Derived from Image Analysis
2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
Junichi Yamagishi - One of the best experts on this subject based on the ideXlab platform.
-
evaluation of speaker verification security and detection of hmm based Synthetic Speech
IEEE Transactions on Audio Speech and Language Processing, 2012Co-Authors: Phillip L. De Leon, Inma Hernaez, Michael Pucher, Junichi Yamagishi, Ibon SaratxagaAbstract:In this paper, we evaluate the vulnerability of speaker verification (SV) systems to Synthetic Speech. The SV systems are based on either the Gaussian mixture model–universal background model (GMM-UBM) or support vector machine (SVM) using GMM supervectors. We use a hidden Markov model (HMM)-based text-to-Speech (TTS) synthesizer, which can synthesize Speech for a target speaker using small amounts of training data through model adaptation of an average voice or background model. Although the SV systems have a very low equal error rate (EER), when tested with Synthetic Speech generated from speaker models derived from the Wall Street Journal (WSJ) Speech corpus, over 81% of the matched claims are accepted. This result suggests vulnerability in SV systems and thus a need to accurately detect Synthetic Speech. We propose a new feature based on relative phase shift (RPS), demonstrate reliable detection of Synthetic Speech, and show how this classifier can be used to improve security of SV systems.
-
Synthetic Speech discrimination using pitch pattern statistics derived from image analysis
Conference of the International Speech Communication Association, 2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
-
INTERSpeech - Synthetic Speech Discrimination using Pitch Pattern Statistics Derived from Image Analysis
2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
-
detection of Synthetic Speech for the problem of imposture
International Conference on Acoustics Speech and Signal Processing, 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to Synthetic Speech. We use a HMM-based Speech synthesizer, which creates Synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with Synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect Synthetic Speech and develop a GMM-based Synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and Synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
-
ICASSP - Detection of Synthetic Speech for the problem of imposture
2011 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to Synthetic Speech. We use a HMM-based Speech synthesizer, which creates Synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with Synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect Synthetic Speech and develop a GMM-based Synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and Synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
Bryan Stewart - One of the best experts on this subject based on the ideXlab platform.
-
ICASSP - Performance of I-vector speaker verification and the detection of Synthetic Speech
2014 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2014Co-Authors: Richard D. Mcclanahan, Bryan Stewart, Phillip L. De LeonAbstract:In this paper, we present new research results on the vulnerability of speaker verification (SV) systems to Synthetic Speech. Using a state-of-the-art i-vector SV system and evaluating with the Wall-Street Journal (WSJ) corpus, our SV system has a 0.00% false rejection rate (FRR) and 1.74 × 10 -5 false acceptance rate (FAR). When the i-vector system is tested with state-of-the-art speaker-adaptive, hidden Markov model (HMM)-based Synthetic Speech generated from speaker models derived from the WSJ journal corpus, 22.9% of the matched claims are accepted highlighting the vulnerability of SV systems to Synthetic Speech. We propose a new Synthetic Speech detector (SSD) which uses previously-proposed features derived from image analysis of pitch patterns but extracted on phoneme-level segments and which leverages the available enrollment Speech from the SV system. When the SSD is applied to human and Synthetic Speech accepted by the SV system, the overall system has a FRR of 7.35% and a FAR of 2.34 × 10 -4 which is lower than previously-reported systems and thus significantly reduces the vulnerability.
-
ICASSP - Synthetic Speech detection based on selectedword discriminators
2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013Co-Authors: Phillip L. De Leon, Bryan StewartAbstract:Speaker verification (SV) systems have been shown to be vulnerable to imposture using Speech synthesizers. In this paper, we extend previous work in detecting Synthetic Speech by analyzing words which provide strong discrimination between human and Synthetic Speech. The research is applicable to authentication systems based on text-dependent SV where the user is prompted to speak a certain utterance which can be chosen by the designer. Our results show that this approach to Synthetic Speech detection leads to higher accuracies than other proposed approaches. Using various corpora to train and test, our results show 98% accuracy in correctly classifying both human and Synthetic Speech.
-
Synthetic Speech discrimination using pitch pattern statistics derived from image analysis
Conference of the International Speech Communication Association, 2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
-
INTERSpeech - Synthetic Speech Discrimination using Pitch Pattern Statistics Derived from Image Analysis
2012Co-Authors: Phillip L. De Leon, Bryan Stewart, Junichi YamagishiAbstract:In this paper, we extend the work by Ogihara, et al. to discriminate between human and Synthetic Speech using features based on pitch patterns. As previously demonstrated, significant differences in pitch patterns between human and Synthetic Speech can be leveraged to classify Speech as being human or Synthetic in origin. We propose using mean pitch stability, mean pitch stability range, and jitter as features extracted after image analysis of pitch patterns. We have observed that for Synthetic Speech, these features lie in a small and distinct space as compared to human Speech and have modeled them with a multivariate Gaussian distribution. Our classifier is trained using Synthetic Speech collected from the 2008 and 2011 Blizzard Challenge along with Festival pre-built voices and human Speech from the NIST2002 corpus. We evaluate the classifier on a much larger corpus than previously studied using human Speech from the Switchboard corpus, Synthetic Speech from the Resource Management corpus, and Synthetic Speech generated from Festival trained on the Wall Street Journal corpus. Results show 98% accuracy in correctly classifying human Speech and 96% accuracy in correctly classifying Synthetic Speech. Index Terms: Speaker recognition, Speech synthesis, Security
Ibon Saratxaga - One of the best experts on this subject based on the ideXlab platform.
-
Synthetic Speech detection using phase information
Speech Communication, 2016Co-Authors: Ibon Saratxaga, Inma Hernaez, Jon Sanchez, Eva NavasAbstract:Phase information based Synthetic Speech detectors (RPS, MGD) are analyzed.Training using real attack samples and copy-synthesized material is evaluated.Evaluation of the detectors against unknown attacks, including channel effect.Detectors work well for voice conversion and adapted Synthetic Speech impostors. Taking advantage of the fact that most of the Speech processing techniques neglect the phase information, we seek to detect phase perturbations in order to prevent Synthetic impostors attacking Speaker Verification systems. Two Synthetic Speech Detection (SSD) systems that use spectral phase related information are reviewed and evaluated in this work: one based on the Modified Group Delay (MGD), and the other based on the Relative Phase Shift, (RPS). A classical module-based MFCC system is also used as baseline. Different training strategies are proposed and evaluated using both real spoofing samples and copy-synthesized signals from the natural ones, aiming to alleviate the issue of getting real data to train the systems. The recently published ASVSpoof2015 database is used for training and evaluation. Performance with completely unrelated data is also checked using Synthetic Speech from the Blizzard Challenge as evaluation material. The results prove that phase information can be successfully used for the SSD task even with unknown attacks.
-
evaluation of speaker verification security and detection of hmm based Synthetic Speech
IEEE Transactions on Audio Speech and Language Processing, 2012Co-Authors: Phillip L. De Leon, Inma Hernaez, Michael Pucher, Junichi Yamagishi, Ibon SaratxagaAbstract:In this paper, we evaluate the vulnerability of speaker verification (SV) systems to Synthetic Speech. The SV systems are based on either the Gaussian mixture model–universal background model (GMM-UBM) or support vector machine (SVM) using GMM supervectors. We use a hidden Markov model (HMM)-based text-to-Speech (TTS) synthesizer, which can synthesize Speech for a target speaker using small amounts of training data through model adaptation of an average voice or background model. Although the SV systems have a very low equal error rate (EER), when tested with Synthetic Speech generated from speaker models derived from the Wall Street Journal (WSJ) Speech corpus, over 81% of the matched claims are accepted. This result suggests vulnerability in SV systems and thus a need to accurately detect Synthetic Speech. We propose a new feature based on relative phase shift (RPS), demonstrate reliable detection of Synthetic Speech, and show how this classifier can be used to improve security of SV systems.
-
detection of Synthetic Speech for the problem of imposture
International Conference on Acoustics Speech and Signal Processing, 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to Synthetic Speech. We use a HMM-based Speech synthesizer, which creates Synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with Synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect Synthetic Speech and develop a GMM-based Synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and Synthetic Speech in 88% of tests thus significantly reducing the vulnerability.
-
ICASSP - Detection of Synthetic Speech for the problem of imposture
2011 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2011Co-Authors: Phillip L. De Leon, Inma Hernaez, Ibon Saratxaga, Michael Pucher, Junichi YamagishiAbstract:In this paper, we present new results from our research into the vulnerability of a speaker verification (SV) system to Synthetic Speech. We use a HMM-based Speech synthesizer, which creates Synthetic Speech for a targeted speaker through adaptation of a background model and both GMM-UBM and support vector machine (SVM) SV systems. Using 283 speakers from the Wall-Street Journal (WSJ) corpus, our SV systems have a 0.35% EER. When the systems are tested with Synthetic Speech generated from speaker models derived from the WSJ journal corpus, over 91% of the matched claims are accepted. We propose the use of relative phase shift (RPS) in order to detect Synthetic Speech and develop a GMM-based Synthetic Speech classifier (SSC). Using the SSC, we are able to correctly classify human Speech in 95% of tests and Synthetic Speech in 88% of tests thus significantly reducing the vulnerability.