The Experts below are selected from a list of 129 Experts worldwide ranked by ideXlab platform

Christian M??ller - One of the best experts on this subject based on the ideXlab platform.

  • Multilingual Speaker age recognition: Regression analyses on the Lwazi corpus
    Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition and Understanding ASRU 2009, 2009
    Co-Authors: Michael Feld, Charl Van Heerden, Etienne Barnard, Christian M??ller
    Abstract:

    Multilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas Multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of Multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely Speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true multi-lingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the Speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors.

Michael Feld - One of the best experts on this subject based on the ideXlab platform.

  • ASRU - Multilingual Speaker age recognition: Regression analyses on the Lwazi corpus
    2009 IEEE Workshop on Automatic Speech Recognition & Understanding, 2009
    Co-Authors: Michael Feld, Etienne Barnard, Charl Van Heerden, Christian Müller
    Abstract:

    Multilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas Multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of Multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely Speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true Multilingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the Speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors.

  • Multilingual Speaker age recognition: Regression analyses on the Lwazi corpus
    Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition and Understanding ASRU 2009, 2009
    Co-Authors: Michael Feld, Charl Van Heerden, Etienne Barnard, Christian M??ller
    Abstract:

    Multilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas Multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of Multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely Speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true multi-lingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the Speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors.

  • towards a Multilingual approach on Speaker classication
    2006
    Co-Authors: Michael Feld
    Abstract:

    This paper outlines a framework for a Multilingual Speaker classication system which is based on an underlying language identication module. First, the AGENDER Speaker classication technology is introduced, a two-layered approach which primarily recognizes the Speakers’ age and gender but also incorporates novel domain-independent aspects that can be applied to other Speaker characteristics like emotions or cognitive load. Then, it is pointed out that one of its major drawbacks consists of the fact that it has not been veried that the chosen set of speech features also works for other languages, especially for those with different phonological aspects. To overcome this drawback, it is suggested to extend AGENDER with a language identication module. The module presented here is designed to meet the requirements of a specic telephone-based application (which itself is not within the focus of this paper): The languages German, English and Turkish shall be discriminated on the basis of the initial utterance of the Speaker; for each of the possible languages, hypotheses about the nature of the initial utterance are available; the domain encompasses a list of English product names. Although the suggested method is as yet only partly implemented, the rst evaluation results are very promising: Turkish could be identied with an accuracy of 71.75 %, German with an accuracy of 78.39 %, and English with an accuracy of 79.89 %. Besides this, the paper outlines the use of the language identication module within a Multilingual version of AGENDER.

H S Jayanna - One of the best experts on this subject based on the ideXlab platform.

  • An Experimental Comparison of Modeling Techniques and Combination of Speaker – Specific Information from Different Languages for Multilingual Speaker Identification
    Journal of intelligent systems, 2016
    Co-Authors: H S Jayanna, B G Nagaraja
    Abstract:

    AbstractMost of the state-of-the-art Speaker identification systems work on a monolingual (preferably English) scenario. Therefore, English-language autocratic countries can use the system efficiently for Speaker recognition. However, there are many countries, including India, that are Multilingual in nature. People in such countries have habituated to speak multiple languages. The existing Speaker identification system may yield poor performance if a Speaker’s train and test data are in different languages. Thus, developing a robust Multilingual Speaker identification system is an issue in many countries. In this work, an experimental evaluation of the modeling techniques, including self-organizing map (SOM), learning vector quantization (LVQ), and Gaussian mixture model-universal background model (GMM-UBM) classifiers for Multilingual Speaker identification, is presented. The monolingual and crosslingual Speaker identification studies are conducted using 50 Speakers of our own database. It is observed from the experimental results that the GMM-UBM classifier gives better identification performance than the SOM and LVQ classifiers. Furthermore, we propose a combination of Speaker-specific information from different languages for crosslingual Speaker identification, and it is observed that the combination feature gives better performance in all the crosslingual Speaker identification experiments.

  • feature extraction and modelling techniques for Multilingual Speaker recognition a review
    International Journal of Signal and Imaging Systems Engineering, 2016
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    In this paper, recognising the Speaker in the Multilingual context is demonstrated. Multilingual Speaker recognition is useful in countries like India, Canada and South Africa where multiple languages are used for communication. In such countries, use of the monolingual Speaker recognition system may not yield the expected level of performance. This is because a Speaker may change his training and test languages whenever it is necessary. As a result, the system yields poor performance due to variation in language parameters. Therefore, developing a robust Multilingual Speaker recognition system is an issue in many countries. Researchers in the field of Speaker recognition have made a few attempts to identify the Speaker in a Multilingual context. This paper gives an overview of various techniques developed for monolingual, cross-lingual and Multilingual Speaker recognition systems.

  • Multilingual Speaker Identification by Combining Evidence from LPR and Multitaper MFCC
    Journal of intelligent systems, 2013
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    AbstractIn this work, the significance of combining the evidence from multitaper mel-frequency cepstral coefficients (MFCC), linear prediction residual (LPR), and linear prediction residual phase (LPRP) features for Multilingual Speaker identification with the constraint of limited data condition is demonstrated. The LPR is derived from linear prediction analysis, and LPRP is obtained by dividing the LPR using its Hilbert envelope. The sine-weighted cepstrum estimators (SWCE) with six tapers are considered for multitaper MFCC feature extraction. The Gaussian mixture model–universal background model is used for modeling each Speaker for different evidence. The evidence is then combined at scoring level to improve the performance. The monolingual, crosslingual, and Multilingual Speaker identification studies were conducted using 30 randomly selected Speakers from the IITG multivariability Speaker recognition database. The experimental results show that the combined evidence improves the performance by nearly 8–10% compared with individual evidence.

  • Multilingual Speaker identification with the constraint of limited data using multitaper mfcc
    International Conference on Security in Computer Networks and Distributed Systems, 2012
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    Feature extraction has the ability to improve the performance of Speaker identification systems. This paper studies the significance of low-variance multitaper Mel-frequency cepstral coefficient (multitaper MFCC) features for Multilingual Speaker identification with the constraint of limited data. The Speaker identification study is conducted using 30 Speakers of our own database. Sine-weighted cepstrum estimator (SWCE) taper MFCC features are extracted and modeled using Gaussian Mixture Model (GMM)-Universal Background Model (UBM). The results show that the multitaper MFCC approach performs better than the conventional Hamming window MFCC technique in all the Speaker identification experiments.

  • SNDS - Multilingual Speaker Identification with the Constraint of Limited Data Using Multitaper MFCC
    Communications in Computer and Information Science, 2012
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    Feature extraction has the ability to improve the performance of Speaker identification systems. This paper studies the significance of low-variance multitaper Mel-frequency cepstral coefficient (multitaper MFCC) features for Multilingual Speaker identification with the constraint of limited data. The Speaker identification study is conducted using 30 Speakers of our own database. Sine-weighted cepstrum estimator (SWCE) taper MFCC features are extracted and modeled using Gaussian Mixture Model (GMM)-Universal Background Model (UBM). The results show that the multitaper MFCC approach performs better than the conventional Hamming window MFCC technique in all the Speaker identification experiments.

Etienne Barnard - One of the best experts on this subject based on the ideXlab platform.

  • ASRU - Multilingual Speaker age recognition: Regression analyses on the Lwazi corpus
    2009 IEEE Workshop on Automatic Speech Recognition & Understanding, 2009
    Co-Authors: Michael Feld, Etienne Barnard, Charl Van Heerden, Christian Müller
    Abstract:

    Multilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas Multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of Multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely Speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true Multilingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the Speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors.

  • Multilingual Speaker age recognition: Regression analyses on the Lwazi corpus
    Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition and Understanding ASRU 2009, 2009
    Co-Authors: Michael Feld, Charl Van Heerden, Etienne Barnard, Christian M??ller
    Abstract:

    Multilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas Multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of Multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely Speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true multi-lingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the Speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors.

  • language dependence in Multilingual Speaker verification
    2005
    Co-Authors: Neil Kleynhans, Etienne Barnard
    Abstract:

    Sixteenth Annual Symposium of the Pattern Recognition Association of South Africa, Langebaan, South Africa, 23-25 November 2005

B G Nagaraja - One of the best experts on this subject based on the ideXlab platform.

  • An Experimental Comparison of Modeling Techniques and Combination of Speaker – Specific Information from Different Languages for Multilingual Speaker Identification
    Journal of intelligent systems, 2016
    Co-Authors: H S Jayanna, B G Nagaraja
    Abstract:

    AbstractMost of the state-of-the-art Speaker identification systems work on a monolingual (preferably English) scenario. Therefore, English-language autocratic countries can use the system efficiently for Speaker recognition. However, there are many countries, including India, that are Multilingual in nature. People in such countries have habituated to speak multiple languages. The existing Speaker identification system may yield poor performance if a Speaker’s train and test data are in different languages. Thus, developing a robust Multilingual Speaker identification system is an issue in many countries. In this work, an experimental evaluation of the modeling techniques, including self-organizing map (SOM), learning vector quantization (LVQ), and Gaussian mixture model-universal background model (GMM-UBM) classifiers for Multilingual Speaker identification, is presented. The monolingual and crosslingual Speaker identification studies are conducted using 50 Speakers of our own database. It is observed from the experimental results that the GMM-UBM classifier gives better identification performance than the SOM and LVQ classifiers. Furthermore, we propose a combination of Speaker-specific information from different languages for crosslingual Speaker identification, and it is observed that the combination feature gives better performance in all the crosslingual Speaker identification experiments.

  • feature extraction and modelling techniques for Multilingual Speaker recognition a review
    International Journal of Signal and Imaging Systems Engineering, 2016
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    In this paper, recognising the Speaker in the Multilingual context is demonstrated. Multilingual Speaker recognition is useful in countries like India, Canada and South Africa where multiple languages are used for communication. In such countries, use of the monolingual Speaker recognition system may not yield the expected level of performance. This is because a Speaker may change his training and test languages whenever it is necessary. As a result, the system yields poor performance due to variation in language parameters. Therefore, developing a robust Multilingual Speaker recognition system is an issue in many countries. Researchers in the field of Speaker recognition have made a few attempts to identify the Speaker in a Multilingual context. This paper gives an overview of various techniques developed for monolingual, cross-lingual and Multilingual Speaker recognition systems.

  • Multilingual Speaker Identification by Combining Evidence from LPR and Multitaper MFCC
    Journal of intelligent systems, 2013
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    AbstractIn this work, the significance of combining the evidence from multitaper mel-frequency cepstral coefficients (MFCC), linear prediction residual (LPR), and linear prediction residual phase (LPRP) features for Multilingual Speaker identification with the constraint of limited data condition is demonstrated. The LPR is derived from linear prediction analysis, and LPRP is obtained by dividing the LPR using its Hilbert envelope. The sine-weighted cepstrum estimators (SWCE) with six tapers are considered for multitaper MFCC feature extraction. The Gaussian mixture model–universal background model is used for modeling each Speaker for different evidence. The evidence is then combined at scoring level to improve the performance. The monolingual, crosslingual, and Multilingual Speaker identification studies were conducted using 30 randomly selected Speakers from the IITG multivariability Speaker recognition database. The experimental results show that the combined evidence improves the performance by nearly 8–10% compared with individual evidence.

  • Multilingual Speaker identification with the constraint of limited data using multitaper mfcc
    International Conference on Security in Computer Networks and Distributed Systems, 2012
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    Feature extraction has the ability to improve the performance of Speaker identification systems. This paper studies the significance of low-variance multitaper Mel-frequency cepstral coefficient (multitaper MFCC) features for Multilingual Speaker identification with the constraint of limited data. The Speaker identification study is conducted using 30 Speakers of our own database. Sine-weighted cepstrum estimator (SWCE) taper MFCC features are extracted and modeled using Gaussian Mixture Model (GMM)-Universal Background Model (UBM). The results show that the multitaper MFCC approach performs better than the conventional Hamming window MFCC technique in all the Speaker identification experiments.

  • SNDS - Multilingual Speaker Identification with the Constraint of Limited Data Using Multitaper MFCC
    Communications in Computer and Information Science, 2012
    Co-Authors: B G Nagaraja, H S Jayanna
    Abstract:

    Feature extraction has the ability to improve the performance of Speaker identification systems. This paper studies the significance of low-variance multitaper Mel-frequency cepstral coefficient (multitaper MFCC) features for Multilingual Speaker identification with the constraint of limited data. The Speaker identification study is conducted using 30 Speakers of our own database. Sine-weighted cepstrum estimator (SWCE) taper MFCC features are extracted and modeled using Gaussian Mixture Model (GMM)-Universal Background Model (UBM). The results show that the multitaper MFCC approach performs better than the conventional Hamming window MFCC technique in all the Speaker identification experiments.