The Experts below are selected from a list of 122694 Experts worldwide ranked by ideXlab platform
Haizhou Li - One of the best experts on this subject based on the ideXlab platform.
-
A Maximum-Entropy Segmentation Model for Statistical Machine Translation
IEEE Transactions on Audio Speech and Language Processing, 2011Co-Authors: Deyi Xiong, Min Zhang, Haizhou LiAbstract:Segmentation is of great importance to statistical machine translation. It splits a source sentence into sequences of translatable segments. We propose a maximum-entropy Segmentation Model to capture desirable phrasal and hierarchical Segmentations for statistical machine translation. We present an approach to automatically learning the beginning and ending boundaries of cohesive segments from word-aligned bilingual data without using any additional resources. The learned boundaries are then used to define cohesive segments in both phrasal and hierarchical Segmentations. We integrate the Segmentation Model into phrasal statistical machine translation (SMT) and conduct experiments on the newswire and broadcast news domain to investigate the effectiveness of the proposed Segmentation Model on a large-scale training data. Our experimental results show that the maximum-entropy Segmentation Model significantly improves translation quality in terms of BLEU. We further validate that 1) the proposed Segmentation Model significantly outperforms syntactic constraints which are used in previous work to constrain Segmentations; and 2) it is necessary to capture hierarchical Segmentations besides phrasal Segmentations.
Christoph Tillmann - One of the best experts on this subject based on the ideXlab platform.
-
a unigram orientation Model for statistical machine translation
North American Chapter of the Association for Computational Linguistics, 2004Co-Authors: Christoph TillmannAbstract:In this paper, we present a unigram Segmentation Model for statistical machine translation where the Segmentation units are blocks: pairs of phrases without internal structure. The Segmentation Model uses a novel orientation component to handle swapping of neighbor blocks. During training, we collect block unigram counts with orientation: we count how often a block occurs to the left or to the right of some predecessor block. The orientation Model is shown to improve translation performance over two Models: 1) no block re-ordering is used, and 2) the block swapping is controlled only by a language Model. We show experimental results on a standard Arabic-English translation task.
Deyi Xiong - One of the best experts on this subject based on the ideXlab platform.
-
A Maximum-Entropy Segmentation Model for Statistical Machine Translation
IEEE Transactions on Audio Speech and Language Processing, 2011Co-Authors: Deyi Xiong, Min Zhang, Haizhou LiAbstract:Segmentation is of great importance to statistical machine translation. It splits a source sentence into sequences of translatable segments. We propose a maximum-entropy Segmentation Model to capture desirable phrasal and hierarchical Segmentations for statistical machine translation. We present an approach to automatically learning the beginning and ending boundaries of cohesive segments from word-aligned bilingual data without using any additional resources. The learned boundaries are then used to define cohesive segments in both phrasal and hierarchical Segmentations. We integrate the Segmentation Model into phrasal statistical machine translation (SMT) and conduct experiments on the newswire and broadcast news domain to investigate the effectiveness of the proposed Segmentation Model on a large-scale training data. Our experimental results show that the maximum-entropy Segmentation Model significantly improves translation quality in terms of BLEU. We further validate that 1) the proposed Segmentation Model significantly outperforms syntactic constraints which are used in previous work to constrain Segmentations; and 2) it is necessary to capture hierarchical Segmentations besides phrasal Segmentations.
Paul Damien - One of the best experts on this subject based on the ideXlab platform.
-
a note on ramaswamy chatterjee and cohen s latent joint Segmentation Models
Journal of Marketing Research, 1999Co-Authors: Stephen G. Walker, Paul DamienAbstract:In this note, the authors show that the latent Segmentation Model developed by Ramaswamy, Chatterjee, and Cohen (1996) is nonidentifiable, and therefore, the parameters of interest are nonestimable.
Steven H Cohen - One of the best experts on this subject based on the ideXlab platform.
-
reply to a note on ramaswamy chatterjee and cohen s latent joint Segmentation Models
Journal of Marketing Research, 1999Co-Authors: Venkatram Ramaswamy, Rabikar Chatterjee, Steven H CohenAbstract:Walker and Damien (1999) assert that the latent Segmentation Model developed by Ramaswamy, Chatterjee, and Cohen (1996) is nonidentifiable. In response, the authors show that this assertion is inco...