The Experts below are selected from a list of 4002 Experts worldwide ranked by ideXlab platform

Houda Saadane - One of the best experts on this subject based on the ideXlab platform.

  • using transliteration of proper names from arabic to Latin Script to improve english arabic word alignment
    International Joint Conference on Natural Language Processing, 2013
    Co-Authors: Nasredine Semmar, Houda Saadane
    Abstract:

    Bilingual lexicons of proper names play a vital role in machine translation and cross-language information retrieval. Word alignment approaches are generally used to construct bilingual lexicons automatically from parallel corpora. Aligning proper names is a task particularly difficult when the source and target languages of the parallel corpus do not share a same written Script. We present in this paper a system to transliterate automatically proper names from Arabic to Latin Script, and a tool to align single and compound words from English-Arabic parallel texts. We particularly focus on the impact of using transliteration to improve the performance of the word alignment tool. We have evaluated the word alignment tool integrating transliteration of proper names from Arabic to Latin Script using two methods: A manual evaluation of the alignment quality and an evaluation of the impact of this alignment on the translation quality by using the open source statistical machine translation system Moses. Experiments show that integrating transliteration of proper names into the alignment process improves the Fmeasure of word alignment from 72% to 81% and the translation BLEU score from 20.15% to 20.63%.

  • IJCNLP - Using Transliteration of Proper Names from Arabic to Latin Script to Improve English-Arabic Word Alignment
    2013
    Co-Authors: Nasredine Semmar, Houda Saadane
    Abstract:

    Bilingual lexicons of proper names play a vital role in machine translation and cross-language information retrieval. Word alignment approaches are generally used to construct bilingual lexicons automatically from parallel corpora. Aligning proper names is a task particularly difficult when the source and target languages of the parallel corpus do not share a same written Script. We present in this paper a system to transliterate automatically proper names from Arabic to Latin Script, and a tool to align single and compound words from English-Arabic parallel texts. We particularly focus on the impact of using transliteration to improve the performance of the word alignment tool. We have evaluated the word alignment tool integrating transliteration of proper names from Arabic to Latin Script using two methods: A manual evaluation of the alignment quality and an evaluation of the impact of this alignment on the translation quality by using the open source statistical machine translation system Moses. Experiments show that integrating transliteration of proper names into the alignment process improves the Fmeasure of word alignment from 72% to 81% and the translation BLEU score from 20.15% to 20.63%.

Philippe Grange - One of the best experts on this subject based on the ideXlab platform.

  • training schemes for the transliteration of the balinese Script into the Latin Script on palm leaf manuScript images
    International Conference on Frontiers in Handwriting Recognition, 2018
    Co-Authors: Made Windu Antara Kesiman, Jeanchristophe Burie, Jeanmarc Ogier, Philippe Grange
    Abstract:

    Considering the importance of the contents of the Balinese palm leaf manuScripts, transliteration system has to be developed in order to be able to read easily these manuScripts. The challenge comes from the fact that Balinese Script is a syllabic Script and the mapping between linguistic symbols and images of symbols is not straightforward. In addition, with a very limited training data availability, some adaptations of LSTM in the transliteration training scheme need to be designed, to be analyzed and to be evaluated. This paper contributes in proposing and evaluating some adapted segmentation free training schemes for the transliteration of the Balinese Script into the Latin Script from palm leaf manuScript images. We describe the generated synthetic dataset and the proposed training schemes at two different levels (word level and text line level) to transliterate the real word and text lines from palm leaf manuScript images. For word transliteration, in general, training schemes at word level perform better than training schemes at text line level. As comparison, the segmentation based transliteration method gives a very promising result. For text line transliteration, segmentation based transliteration method outperforms all segmentation free training schemes for the less degraded collections, while the segmentation free training schemes contributes in transliterating the text lines for more degraded manuScripts. Training at text line level with a pre-trained model at word level could give a better result in word transliteration while still keeping the optimal performances for text line transliteration.

  • ICFHR - Training Schemes for the Transliteration of the Balinese Script Into the Latin Script on Palm Leaf ManuScript Images
    2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR), 2018
    Co-Authors: Made Windu Antara Kesiman, Jeanchristophe Burie, Jeanmarc Ogier, Philippe Grange
    Abstract:

    Considering the importance of the contents of the Balinese palm leaf manuScripts, transliteration system has to be developed in order to be able to read easily these manuScripts. The challenge comes from the fact that Balinese Script is a syllabic Script and the mapping between linguistic symbols and images of symbols is not straightforward. In addition, with a very limited training data availability, some adaptations of LSTM in the transliteration training scheme need to be designed, to be analyzed and to be evaluated. This paper contributes in proposing and evaluating some adapted segmentation free training schemes for the transliteration of the Balinese Script into the Latin Script from palm leaf manuScript images. We describe the generated synthetic dataset and the proposed training schemes at two different levels (word level and text line level) to transliterate the real word and text lines from palm leaf manuScript images. For word transliteration, in general, training schemes at word level perform better than training schemes at text line level. As comparison, the segmentation based transliteration method gives a very promising result. For text line transliteration, segmentation based transliteration method outperforms all segmentation free training schemes for the less degraded collections, while the segmentation free training schemes contributes in transliterating the text lines for more degraded manuScripts. Training at text line level with a pre-trained model at word level could give a better result in word transliteration while still keeping the optimal performances for text line transliteration.

Zhang Jian - One of the best experts on this subject based on the ideXlab platform.

  • Chinese-Uyghur machine translation system for phrase-based statistical translation
    Journal of Computer Applications, 2009
    Co-Authors: Zhang Jian
    Abstract:

    This paper described a Chinese-Uyghur machine translation system for phrase-based statistical translation.First the Chinese-Uyghur corpus was used to train the language model and the translation model.After using these models to decode the source statement,the authors got the best translation statement.The core decoding algorithm was the beam search algorithm,and the Latin-Script Uyghur was used as the Uyghur corpus.The experimental results show that phrase-based statistical machine translation method can be used quickly and effectively to build a Chinese-Uygher machine translation platform.

Brian Roark - One of the best experts on this subject based on the ideXlab platform.

  • Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset
    arXiv: Computation and Language, 2020
    Co-Authors: Brian Roark, Lawrence Wolf-sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
    Abstract:

    This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native Scripts for 12 South Asian languages. The dataset includes, for each language: 1) native Script Wikipedia text; 2) a romanization lexicon; and 3) full sentence parallel data in both a native Script of the language and the basic Latin alphabet. We document the methods used for preparation and selection of the Wikipedia text in each language; collection of attested romanizations for sampled lexicons; and manual romanization of held-out sentences from the native Script collections. We additionally provide baseline results on several tasks made possible by the dataset, including single word transliteration, full sentence transliteration, and language modeling of native Script and romanized text. Keywords: romanization, transliteration, South Asian languages

  • processing south asian languages written in the Latin Script the dakshina dataset
    Language Resources and Evaluation, 2020
    Co-Authors: Brian Roark, Lawrence Wolfsonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
    Abstract:

    This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native Scripts for 12 South Asian languages. The dataset includes, for each language: 1) native Script Wikipedia text; 2) a romanization lexicon; and 3) full sentence parallel data in both a native Script of the language and the basic Latin alphabet. We document the methods used for preparation and selection of the Wikipedia text in each language; collection of attested romanizations for sampled lexicons; and manual romanization of held-out sentences from the native Script collections. We additionally provide baseline results on several tasks made possible by the dataset, including single word transliteration, full sentence transliteration, and language modeling of native Script and romanized text.

  • LREC - Processing South Asian languages written in the Latin Script: the Dakshina dataset
    2020
    Co-Authors: Brian Roark, Lawrence Wolf-sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
    Abstract:

    This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native Scripts for 12 South Asian languages. The dataset includes, for each language: 1) native Script Wikipedia text; 2) a romanization lexicon; and 3) full sentence parallel data in both a native Script of the language and the basic Latin alphabet. We document the methods used for preparation and selection of the Wikipedia text in each language; collection of attested romanizations for sampled lexicons; and manual romanization of held-out sentences from the native Script collections. We additionally provide baseline results on several tasks made possible by the dataset, including single word transliteration, full sentence transliteration, and language modeling of native Script and romanized text.

  • Latin Script keyboards for south asian languages with finite state normalization
    Finite-State Methods and Natural Language Processing, 2019
    Co-Authors: Lawrence Wolfsonkin, Vlad Schogol, Brian Roark, Michael A Riley
    Abstract:

    The use of the Latin Script for text entry of South Asian languages is common, even though there is no standard orthography for these languages in the Script. We explore several compact finite-state architectures that permit variable spellings of words during mobile text entry. We find that approaches making use of transliteration transducers provide large accuracy improvements over baselines, but that simpler approaches involving a compact representation of many attested alternatives yields much of the accuracy gain. This is particularly important when operating under constraints on model size (e.g., on inexpensive mobile devices with limited storage and memory for keyboard models), and on speed of inference, since people typing on mobile keyboards expect no perceptual delay in keyboard responsiveness.

  • FSMNLP - Latin Script keyboards for South Asian languages with finite-state normalization
    Proceedings of the 14th International Conference on Finite-State Methods and Natural Language Processing, 2019
    Co-Authors: Lawrence Wolf-sonkin, Vlad Schogol, Brian Roark, Michael A Riley
    Abstract:

    The use of the Latin Script for text entry of South Asian languages is common, even though there is no standard orthography for these languages in the Script. We explore several compact finite-state architectures that permit variable spellings of words during mobile text entry. We find that approaches making use of transliteration transducers provide large accuracy improvements over baselines, but that simpler approaches involving a compact representation of many attested alternatives yields much of the accuracy gain. This is particularly important when operating under constraints on model size (e.g., on inexpensive mobile devices with limited storage and memory for keyboard models), and on speed of inference, since people typing on mobile keyboards expect no perceptual delay in keyboard responsiveness.

Nicole Vincent - One of the best experts on this subject based on the ideXlab platform.

  • icdar2017 competition on the classification of medieval handwritings in Latin Script
    International Conference on Document Analysis and Recognition, 2017
    Co-Authors: Florence Cloppet, Véronique Eglin, Nicole Vincent, Marlene Heliasbaron, Cuong Kieu, Dominique Stutzmann
    Abstract:

    This paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuScripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given.

  • ICDAR 2017 Competition on the Classification of Medieval Handwritings in Latin Script
    2017
    Co-Authors: Florence Cloppet, Véronique Eglin, Van Kieu, Dominique Stutzmann, Marlene Helias-baron, Nicole Vincent
    Abstract:

    This paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuScripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given.

  • ICDAR - ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script
    2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), 2017
    Co-Authors: Florence Cloppet, Véronique Eglin, Nicole Vincent, Cuong Kieu, Marlene Helias-baron, Dominique Stutzmann
    Abstract:

    This paper presents the results of the ICDAR2017 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work follows a competition at ICFHR2016 and aims at providing a rich annotated database of European medieval manuScripts to the community on Handwriting Analysis and Recognition. We proposed four independent classification tasks which attracted 10 registered teams, with 6 submitted classifiers from 4 participants. Those classifiers are trained on a set of 3540 images with their ground truths. In task 1 (Script classification) and task 3 (Date classification), the classifiers have been evaluated by a test set of 2000 greyscale, tiff, 300 dpi images. In task 2 (Script classification) and task 4 (Date classification), the test set consists of 1000 images in different formats, resolutions and color representation. The best scores are respectively 85.2% for task 1, 76.5% for task 2, 59% for task 3, and 49.9% for task 4. An analysis based on the matrix of confusion of each classifier is also given.

  • ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script
    2016
    Co-Authors: Florence Cloppet, Véronique Eglin, Van Kieu, Dominique Stutzmann, Nicole Vincent
    Abstract:

    This paper presents the results of the ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work aims at providing a rich database of European medieval manuScripts to the community on Handwriting Analysis and Recognition. At this competition, we proposed two independent classification tasks which attracted five participants with seven submitted classifiers. Those classifiers are trained on a set 2000 images with their ground truths. In the first task – Script classification – the classifiers have been evaluated by a test set of 1000 single-type manuScripts. In the second task, a ―Fuzzy Classification‖ has been carried out on a set of 2000 multi-Script-type manuScripts. The results of the participants provide the first baseline evaluation up to the accuracy score of 83.9% for the task 1 and to the fuzzy weighted score of 2.96/4 for the task 2. An analysis based on the intra-class distance and matrix of confusion of each classifier is also given.

  • icfhr2016 competition on the classification of medieval handwritings in Latin Script
    International Conference on Frontiers in Handwriting Recognition, 2016
    Co-Authors: Florence Cloppet, Véronique Eglin, Van Kieu, Dominique Stutzmann, Nicole Vincent
    Abstract:

    This paper presents the results of the ICFHR2016 Competition on the Classification of Medieval Handwritings in Latin Script (CLaMM), jointly organized by Computer Scientists and Humanists (paleographers). This work aims at providing a rich database of European medieval manuScripts to the community on Handwriting Analysis and Recognition. At this competition, we proposed two independent classification tasks which attracted five participants with seven submitted classifiers. Those classifiers are trained on a set of 2000 images with their ground truths. In the first task of Script crisp classification, the classifiers have been evaluated on a test set of 1000 single-type manuScripts. In the second task of "Fuzzy Classification", the classifiers have been carried out on a set of 2000 multi-Script-type manuScripts. The results of the participants provide the first baseline evaluation up to the accuracy score of 83.9% for the task 1 and to the fuzzy weighted score of 2.96/4 for the task 2. An analysis based on the intra-class distance and matrix of confusion of each classifier is also given.