The Experts below are selected from a list of 26514 Experts worldwide ranked by ideXlab platform

Ralf Schlüter - One of the best experts on this subject based on the ideXlab platform.

  • Investigations on the use of Morpheme level features in Language Models for Arabic LVCSR
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Amr El-desoky Mousa, Ralf Schlüter
    Abstract:

    A major challenge for Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) is the rich morphology of Arabic, which leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, the use of Morphemes rather than full-words is considered a better choice for LMs. Thereby, higher lexical coverage and less LM perplexities are achieved. On the other side, an effective way to increase the robustness of LMs is to incorporate features of words into LMs. In this paper, we investigate the use of features derived for Morphemes rather than words. Thus, we combine the benefits of both Morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and Factored LMs (FLMs) estimated over sequences of Morphemes and their features for performing Arabic LVCSR. A relative reduction of 3.9% in Word Error Rate (WER) is achieved compared to a word-based system.

  • Morpheme level feature based language models for german lvcsr
    Conference of the International Speech Communication Association, 2012
    Co-Authors: Amr El-desoky Mousa, Ali Basha M Shaik, Ralf Schlüter
    Abstract:

    One of the challenges for Large Vocabulary Continuous Speech Recognition (LVCSR) of German is its complex morphology and high level of compounding. It leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, building LMs on Morpheme level can be considered a better choice. Thereby, higher lexical coverage and lower LM perplexities are achieved. On the other side, a successful approach to improve the LM probability estimation is to incorporate features of words using feature-based LMs. In this paper, we use features derived for Morphemes as well as words. Thus, we combine the benefits of both Morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and factored LMs (FLMs). Relative reductions of around 1.5% in Word Error Rate (WER) are achieved compared to the best previous results obtained using FLMs.

Amr El-desoky Mousa - One of the best experts on this subject based on the ideXlab platform.

  • Investigations on the use of Morpheme level features in Language Models for Arabic LVCSR
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Amr El-desoky Mousa, Ralf Schlüter
    Abstract:

    A major challenge for Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) is the rich morphology of Arabic, which leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, the use of Morphemes rather than full-words is considered a better choice for LMs. Thereby, higher lexical coverage and less LM perplexities are achieved. On the other side, an effective way to increase the robustness of LMs is to incorporate features of words into LMs. In this paper, we investigate the use of features derived for Morphemes rather than words. Thus, we combine the benefits of both Morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and Factored LMs (FLMs) estimated over sequences of Morphemes and their features for performing Arabic LVCSR. A relative reduction of 3.9% in Word Error Rate (WER) is achieved compared to a word-based system.

  • Morpheme level feature based language models for german lvcsr
    Conference of the International Speech Communication Association, 2012
    Co-Authors: Amr El-desoky Mousa, Ali Basha M Shaik, Ralf Schlüter
    Abstract:

    One of the challenges for Large Vocabulary Continuous Speech Recognition (LVCSR) of German is its complex morphology and high level of compounding. It leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities. In such cases, building LMs on Morpheme level can be considered a better choice. Thereby, higher lexical coverage and lower LM perplexities are achieved. On the other side, a successful approach to improve the LM probability estimation is to incorporate features of words using feature-based LMs. In this paper, we use features derived for Morphemes as well as words. Thus, we combine the benefits of both Morpheme level and feature rich modeling. We compare the performance of stream-based, class-based and factored LMs (FLMs). Relative reductions of around 1.5% in Word Error Rate (WER) are achieved compared to the best previous results obtained using FLMs.

  • Morpheme based factored language models for german lvcsr
    Conference of the International Speech Communication Association, 2011
    Co-Authors: Amr El-desoky Mousa, Basha Shaik, Ralf Schl
    Abstract:

    German is a highly inflectional language, where a large number of words can be generated from the same root. It makes a liberal use of compounding leading to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probability estimates. Therefore, the use of Morphemes for language modeling is considered a better choice for Large Vocabulary Continuous Speech Recognition (LVCSR) than the full-words. Thereby, better lexical coverage and less LM perplexities are achieved. On the other side, the use of Factored Language Models (FLMs) is considered a successful approach that allows the integration of many information sources to get better LM probability estimates. In this paper, we try a combined methodology for language modeling where both morphological decomposition and factored language modeling are used in one model called Morpheme based FLM. Finally, we obtain around 2.5% relative reduction in Word Error Rate (WER) with respect to a traditional full-words system. Index Terms: Morpheme, factored language model, German

Anoop Sarkar - One of the best experts on this subject based on the ideXlab platform.

  • combining Morpheme based machine translation with post processing Morpheme prediction
    Meeting of the Association for Computational Linguistics, 2011
    Co-Authors: Ann Clifton, Anoop Sarkar
    Abstract:

    This paper extends the training and tuning regime for phrase-based statistical machine translation to obtain fluent translations into morphologically complex languages (we build an English to Finnish translation system). Our methods use unsupervised morphology induction. Unlike previous work we focus on morphologically productive phrase pairs -- our decoder can combine Morphemes across phrase boundaries. Morphemes in the target language may not have a corresponding Morpheme or word in the source language. Therefore, we propose a novel combination of post-processing morphology prediction with Morpheme-based translation. We show, using both automatic evaluation scores and linguistically motivated analyses of the output, that our methods outperform previously proposed ones and provide the best known results on the English-Finnish Europarl translation task. Our methods are mostly language independent, so they should improve translation into other target languages with complex morphology.

  • ACL - Combining Morpheme-based Machine Translation with Post-processing Morpheme Prediction
    2011
    Co-Authors: Ann Clifton, Anoop Sarkar
    Abstract:

    This paper extends the training and tuning regime for phrase-based statistical machine translation to obtain fluent translations into morphologically complex languages (we build an English to Finnish translation system). Our methods use unsupervised morphology induction. Unlike previous work we focus on morphologically productive phrase pairs -- our decoder can combine Morphemes across phrase boundaries. Morphemes in the target language may not have a corresponding Morpheme or word in the source language. Therefore, we propose a novel combination of post-processing morphology prediction with Morpheme-based translation. We show, using both automatic evaluation scores and linguistically motivated analyses of the output, that our methods outperform previously proposed ones and provide the best known results on the English-Finnish Europarl translation task. Our methods are mostly language independent, so they should improve translation into other target languages with complex morphology.

Askar Hamdulla - One of the best experts on this subject based on the ideXlab platform.

  • Discriminative approach to lexical entry selection for automatic speech recognition of agglutinative language
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Mijit Ablimit, Tatsuya Kawahara, Askar Hamdulla
    Abstract:

    In agglutinative languages, selection of lexical unit is not obvious. Morpheme unit is usually adopted to ensure the sufficient coverage, but many Morphemes are short, resulting in weak constraints and possible confusions. In this paper, we propose a discriminative approach to select lexical entries which will directly contribute to ASR error reduction. We define an evaluation function for each word by a set of features and their weights, and the measure for optimization by the difference of WERs by the Morpheme-based model and by the word-based model. Then, the weights of the features are learned by a perceptron algorithm. Finally, word (or sub-word) entries with higher evaluation scores are selected to be added to the lexicon. This method is successfully applied to an Uyghur large-vocabulary continuous speech recognition system, resulting in a significant reduction of WER and the lexicon size. Further improvement is achieved by combining with a statistical method based on mutual information criterion.

William D Marslenwilson - One of the best experts on this subject based on the ideXlab platform.

  • derivational morphology and base Morpheme frequency
    Journal of Memory and Language, 2010
    Co-Authors: Michael A Ford, Matthew H Davis, William D Marslenwilson
    Abstract:

    Morpheme frequency effects for derived words (e.g. an influence of the frequency of the base ‘‘dark” on responses to ‘‘darkness”) have been interpreted as evidence of morphemic representation. However, it has been suggested that most derived words would not show these effects if family size (a type frequency count claimed to reflect semantic relationships between whole forms) were controlled. This study used visual lexical decision experiments with correlational designs to compare the influences of base Morpheme frequency and family size on response times to derived words in English and to test for interactions of these variables with suffix productivity. Multiple regression showed that base Morpheme frequency and family size were independent predictors of response times to derived words. Base Morpheme frequency facilitated responses but only to productively suffixed derived words, whereas family size facilitated responses irrespective of productivity. This suggests that base Morpheme frequency effects are independent of Morpheme family size, depend on suffix productivity and indicate that productively suffixed words are represented as Morphemes.