The Experts below are selected from a list of 3690 Experts worldwide ranked by ideXlab platform
Keiichi Tokuda - One of the best experts on this subject based on the ideXlab platform.
-
SSW - The HMM-based Speech Synthesis System (HTS) version 2.0.
2020Co-Authors: Takashi Nose, Takashi Masuko, Junichi Yamagishi, Shinji Sako, Alan W Black, Keiichi TokudaAbstract:A statistical parametric Speech Synthesis System based on hidden Markov models (HMMs) has grown in popularity over the last few years. This System simultaneously models spectrum, excitation, and duration of Speech using context-dependent HMMs and generates Speech waveforms from the HMMs themselves. Since December 2002, we have publicly released an open-source software toolkit named HMM-based Speech Synthesis System (HTS) to provide a research and development platform for the Speech Synthesis community. In December 2006, HTS version 2.0 was released. This version includes a number of new features which are useful for both Speech Synthesis researchers and developers. This paper describes HTS version 2.0 in detail, as well as future release plans.
-
Overview of NIT HMM-based Speech Synthesis System for Blizzard Challenge 2010
2020Co-Authors: Keiichiro Oura, Kei Hashimoto, Sayaka Shiota, Keiichi TokudaAbstract:This paper describes a hidden Markov model (HMM)-based Speech Synthesis System developed for the Blizzard Challenge 2010. This System employs STRAIGHT vocoding, minimum generation error (MGE) training, minimum generation error linear regression (MGELR) based model adaptation, the Bayesian Speech Synthesis framework, and the parameter generation algorithm considering global variance. The real-time factor of the Speech Synthesis System is about 0.3, and its footprint is less than 25 MB. Subjective evaluation results show that the overall Speech quality and intelligibility of the Systems are better than most other System, especially when a well-labeled Speech database can be used. Index Terms: HMM, Speech Synthesis, speaker adaptation, HTS, Blizzard Challenge
-
Overview of NIT HMM-based Speech Synthesis System for Blizzard Challenge 2011
2020Co-Authors: Kei Hashimoto, Keiichiro Oura, Shinji Takaki, Keiichi TokudaAbstract:This paper describes a hidden Markov model (HMM) based Speech Synthesis System developed for the Blizzard Challenge 2011. In the Blizzard Challenge 2011, we focused on the training algorithm for HMM-based Speech Synthesis Systems. To alleviate the local maxima problems in the maximum likelihood estimation, we apply the deterministic annealing expectation maximization (DAEM) algorithm for training HMMs. By using the DAEM algorithm, the reliable acoustic model parameters can be estimated. In addition, we apply stepwise model selection to the model training. The decision tree based context clustering is used as model selection in HMM-based Speech Synthesis. By using the stepwise model selection method, decision trees are gradually changed from small trees into large trees for estimating reliable acoustic models. Subjective evaluation results show that the System synthesized the high intelligible Speech. Index Terms: Speech Synthesis, hidden Markov model, deterministic annealing, model structure
-
recent development of the hmm based Speech Synthesis System hts
Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2009Co-Authors: Keiichiro Oura, Takashi Masuko, Takashi Nose, Junichi Yamagishi, Shinji Sako, Alan W Black, Tomoki Toda, Keiichi TokudaAbstract:A statistical parametric approach to Speech Synthesis based on hidden Markov models (HMMs) has grown in popularity over the last few years. In this approach, spectrum, excitation, and duration of Speech are simultaneously modeled by context-dependent HMMs, and Speech waveforms are generated from the HMMs themselves. Since December 2002, we have publicly released an open-source software toolkit named “HMMbased Speech Synthesis System (HTS)” to provide a research and development toolkit for statistical parametric Speech Synthesis. This paper describes recent developments of HTS in detail, as well as future release plans.
-
the nitech naist hmm based Speech Synthesis System for the blizzard challenge 2006
IEICE Transactions on Information and Systems, 2008Co-Authors: Tomoki Toda, Keiichi TokudaAbstract:We describe a statistical parametric Speech Synthesis System developed by a joint group from the Nagoya Institute of Technology (Nitech) and the Nara Institute of Science and Technology (NAIST) for the annual open evaluation of text-to-Speech Synthesis Systems named Blizzard Challenge 2006. To improve our 2005 System (Nitech-HTS 2005), we investigated new features such as mel-generalized cepstrum-based line spectral pairs (MGC-LSPs), maximum likelihood linear transform (MLLT), and a full covariance global variance (GV) probability density function (pdf). A combination of mel-cepstral coefficients, MLLT, and full covariance GV pdf scored highest in subjective listening tests, and the 2006 System performed significantly better than the 2005 System. The Blizzard Challenge 2006 evaluations show that Nitech-NAIST-HTS 2006 is competitive even when working with relatively large Speech databases.
Tomoki Toda - One of the best experts on this subject based on the ideXlab platform.
-
recent development of the hmm based Speech Synthesis System hts
Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2009Co-Authors: Keiichiro Oura, Takashi Masuko, Takashi Nose, Junichi Yamagishi, Shinji Sako, Alan W Black, Tomoki Toda, Keiichi TokudaAbstract:A statistical parametric approach to Speech Synthesis based on hidden Markov models (HMMs) has grown in popularity over the last few years. In this approach, spectrum, excitation, and duration of Speech are simultaneously modeled by context-dependent HMMs, and Speech waveforms are generated from the HMMs themselves. Since December 2002, we have publicly released an open-source software toolkit named “HMMbased Speech Synthesis System (HTS)” to provide a research and development toolkit for statistical parametric Speech Synthesis. This paper describes recent developments of HTS in detail, as well as future release plans.
-
the nitech naist hmm based Speech Synthesis System for the blizzard challenge 2006
IEICE Transactions on Information and Systems, 2008Co-Authors: Tomoki Toda, Keiichi TokudaAbstract:We describe a statistical parametric Speech Synthesis System developed by a joint group from the Nagoya Institute of Technology (Nitech) and the Nara Institute of Science and Technology (NAIST) for the annual open evaluation of text-to-Speech Synthesis Systems named Blizzard Challenge 2006. To improve our 2005 System (Nitech-HTS 2005), we investigated new features such as mel-generalized cepstrum-based line spectral pairs (MGC-LSPs), maximum likelihood linear transform (MLLT), and a full covariance global variance (GV) probability density function (pdf). A combination of mel-cepstral coefficients, MLLT, and full covariance GV pdf scored highest in subjective listening tests, and the 2006 System performed significantly better than the 2005 System. The Blizzard Challenge 2006 evaluations show that Nitech-NAIST-HTS 2006 is competitive even when working with relatively large Speech databases.
-
performance evaluation of the speaker independent hmm based Speech Synthesis System hts 2007 for the blizzard challenge 2007
International Conference on Acoustics Speech and Signal Processing, 2008Co-Authors: Junichi Yamagishi, Takashi Nose, Tomoki Toda, Keiichi TokudaAbstract:This paper describes a speaker-independent/adaptive HMM-based Speech Synthesis System developed for the Blizzard Challenge 2007. The new System, named "HTS-2007", employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous Systems. Subjective evaluation results show that the new System generates significantly better quality synthetic Speech than that of speaker-dependent approaches with realistic amounts of Speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of Speech data are available.
-
speaker independent hmm based Speech Synthesis System hts 2007 System for the blizzard challenge 2007
2007Co-Authors: Junichi Yamagishi, Tomoki Toda, Keiichi TokudaAbstract:This paper describes an HMM-based Speech Synthesis System developed by the HTS working group for the Blizzard Challenge 2007. To further explore the potential of HMM-based Speech Synthesis, we incorporate new features in our conventional System which underpin a speaker-independent approach: speaker adaptation techniques; adaptive training for HSMMs; and full covariance modeling using the CSMAPLR transforms.
-
details of the nitech hmm based Speech Synthesis System for the blizzard challenge 2005
IEICE Transactions on Information and Systems, 2007Co-Authors: Tomoki Toda, Masaru Nakamura, Keiichi TokudaAbstract:In January 2005, an open evaluation of corpus-based text-to-Speech Synthesis Systems using common Speech datasets, named Blizzard Challenge 2005, was conducted. Nitech group participated in this challenge, entering an HMM-based Speech Synthesis System called Nitech-HTS 2005. This paper describes the technical details, building processes, and performance of our System. We first give an overview of the basic HMM-based Speech Synthesis System, and then describe new features integrated into Nitech-HTS 2005 such as STRAIGHT-based vocoding, HSMM-based acoustic modeling, and a Speech parameter generation algorithm considering GV. Constructed Nitech-HTS 2005 voices can generate Speech waveforms at 0.3 ×RT (real-time ratio) on a 1.6 GHz Pentium 4 machine, and footprints of these voices are less than 2 Mbytes. Subjective listening tests showed that the naturalness and intelligibility of the Nitech-HTS 2005 voices were much better than expected.
Ellen Marie Eide - One of the best experts on this subject based on the ideXlab platform.
-
INTERSpeech - The IBM expressive Speech Synthesis System.
2020Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
SSW - Current status of the IBM Trainable Speech Synthesis System.
2020Co-Authors: Robert E Donovan, Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, Abraham Ittycheriah, Martin Franz, Bhuvana Ramabhadran, Mahesh Viswanathan, Philip GleasonAbstract:This paper describes the current status of the IBM Trainable Speech Synthesis System. The System is a state-of-the-art, trainable, unit-selection based concatenative Speech Synthesiser. The System uses hidden Markov models (HMMs) to provide a phonetic transcription and HMM state alignment of a database of single-speaker continuous-Speech training data. The runtime Synthesiser uses the HMM state sized segments that result as its basic Synthesis units. It determines which segments to concatenate to produce a target sentence using decision trees built from the training data and a dynamic programming search to optimise a perceptually motivated cost function. The Synthesiser can operate both in general domain Text-to-Speech mode, and in Phrase Splicing mode to provide higher quality Synthesis in limited domains. Systems have been built in at least 10 different languages and over 70 voices.
-
the ibm expressive Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
reconciling pronunciation differences between the front end and the back end in the ibm Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo BakisAbstract:In this paper, methods for reconciling pronunciation differences between a rule-based front-end and the pronunciations observed in a database of recorded Speech are presented. The methods are applied to the IBM Expressive Speech Synthesis System [1] for both unrestricted and limited-domain text-to-Speech Synthesis. One method is based on constructing a multiple pronunciation lattice for the given sentence and scoring it using word and phoneme n-gram statistics computed from the target speaker’s database. A second method consists of storing observed pronunciations and introducing them as alternates in the search. We compare the strengths and weaknesses of these two methods. Results show that improvements are achieved in both limited and unrestricted domains, with the largest gains coming in the limited-domain case.
-
current status of the ibm trainable Speech Synthesis System
SSW, 2001Co-Authors: Robert E Donovan, Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, Abraham Ittycheriah, Martin Franz, Bhuvana Ramabhadran, Mahesh Viswanathan, Philip GleasonAbstract:This paper describes the current status of the IBM Trainable Speech Synthesis System. The System is a state-of-the-art, trainable, unit-selection based concatenative Speech Synthesiser. The System uses hidden Markov models (HMMs) to provide a phonetic transcription and HMM state alignment of a database of single-speaker continuous-Speech training data. The runtime Synthesiser uses the HMM state sized segments that result as its basic Synthesis units. It determines which segments to concatenate to produce a target sentence using decision trees built from the training data and a dynamic programming search to optimise a perceptually motivated cost function. The Synthesiser can operate both in general domain Text-to-Speech mode, and in Phrase Splicing mode to provide higher quality Synthesis in limited domains. Systems have been built in at least 10 different languages and over 70 voices.
Wael Hamza - One of the best experts on this subject based on the ideXlab platform.
-
INTERSpeech - The IBM expressive Speech Synthesis System.
2020Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
INTERSpeech - On Building a Concatenative Speech Synthesis System from the Blizzard Challenge Speech Databases
2020Co-Authors: Wael Hamza, Raimo Bakis, Zhi Wei ShuangAbstract:In this paper, we compare two methods of building a concatenative Speech Synthesis System from the relatively small, “Blizzard Challenge” Speech databases. In the first method we build a System directly from the Blizzard databases using the IBM Concatenetative Speech Synthesis System originally designed for very large Speech databases. In the second method, a larger database is used to build the Synthesis System and the output is “morphed” to match the speakers in the Blizzard databases. The second method outperformed the first while maintaining the identity of the Blizzard target speakers.
-
SSW - Current status of the IBM Trainable Speech Synthesis System.
2020Co-Authors: Robert E Donovan, Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, Abraham Ittycheriah, Martin Franz, Bhuvana Ramabhadran, Mahesh Viswanathan, Philip GleasonAbstract:This paper describes the current status of the IBM Trainable Speech Synthesis System. The System is a state-of-the-art, trainable, unit-selection based concatenative Speech Synthesiser. The System uses hidden Markov models (HMMs) to provide a phonetic transcription and HMM state alignment of a database of single-speaker continuous-Speech training data. The runtime Synthesiser uses the HMM state sized segments that result as its basic Synthesis units. It determines which segments to concatenate to produce a target sentence using decision trees built from the training data and a dynamic programming search to optimise a perceptually motivated cost function. The Synthesiser can operate both in general domain Text-to-Speech mode, and in Phrase Splicing mode to provide higher quality Synthesis in limited domains. Systems have been built in at least 10 different languages and over 70 voices.
-
the ibm expressive Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
reconciling pronunciation differences between the front end and the back end in the ibm Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo BakisAbstract:In this paper, methods for reconciling pronunciation differences between a rule-based front-end and the pronunciations observed in a database of recorded Speech are presented. The methods are applied to the IBM Expressive Speech Synthesis System [1] for both unrestricted and limited-domain text-to-Speech Synthesis. One method is based on constructing a multiple pronunciation lattice for the given sentence and scoring it using word and phoneme n-gram statistics computed from the target speaker’s database. A second method consists of storing observed pronunciations and introducing them as alternates in the search. We compare the strengths and weaknesses of these two methods. Results show that improvements are achieved in both limited and unrestricted domains, with the largest gains coming in the limited-domain case.
Raimo Bakis - One of the best experts on this subject based on the ideXlab platform.
-
INTERSpeech - The IBM expressive Speech Synthesis System.
2020Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
INTERSpeech - On Building a Concatenative Speech Synthesis System from the Blizzard Challenge Speech Databases
2020Co-Authors: Wael Hamza, Raimo Bakis, Zhi Wei ShuangAbstract:In this paper, we compare two methods of building a concatenative Speech Synthesis System from the relatively small, “Blizzard Challenge” Speech databases. In the first method we build a System directly from the Blizzard databases using the IBM Concatenetative Speech Synthesis System originally designed for very large Speech databases. In the second method, a larger database is used to build the Synthesis System and the output is “morphed” to match the speakers in the Blizzard databases. The second method outperformed the first while maintaining the identity of the Blizzard target speakers.
-
SSW - Current status of the IBM Trainable Speech Synthesis System.
2020Co-Authors: Robert E Donovan, Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, Abraham Ittycheriah, Martin Franz, Bhuvana Ramabhadran, Mahesh Viswanathan, Philip GleasonAbstract:This paper describes the current status of the IBM Trainable Speech Synthesis System. The System is a state-of-the-art, trainable, unit-selection based concatenative Speech Synthesiser. The System uses hidden Markov models (HMMs) to provide a phonetic transcription and HMM state alignment of a database of single-speaker continuous-Speech training data. The runtime Synthesiser uses the HMM state sized segments that result as its basic Synthesis units. It determines which segments to concatenate to produce a target sentence using decision trees built from the training data and a dynamic programming search to optimise a perceptually motivated cost function. The Synthesiser can operate both in general domain Text-to-Speech mode, and in Phrase Splicing mode to provide higher quality Synthesis in limited domains. Systems have been built in at least 10 different languages and over 70 voices.
-
the ibm expressive Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo Bakis, Michael Picheny, John F PitrelliAbstract:This paper introduces the IBM Expressive Speech Synthesis System. We describe recent work in improving the quality of our baseline text-to-Speech System as well as extending our capabilities to generate expressive synthetic Speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML).
-
reconciling pronunciation differences between the front end and the back end in the ibm Speech Synthesis System
Conference of the International Speech Communication Association, 2004Co-Authors: Wael Hamza, Ellen Marie Eide, Raimo BakisAbstract:In this paper, methods for reconciling pronunciation differences between a rule-based front-end and the pronunciations observed in a database of recorded Speech are presented. The methods are applied to the IBM Expressive Speech Synthesis System [1] for both unrestricted and limited-domain text-to-Speech Synthesis. One method is based on constructing a multiple pronunciation lattice for the given sentence and scoring it using word and phoneme n-gram statistics computed from the target speaker’s database. A second method consists of storing observed pronunciations and introducing them as alternates in the search. We compare the strengths and weaknesses of these two methods. Results show that improvements are achieved in both limited and unrestricted domains, with the largest gains coming in the limited-domain case.