The Experts below are selected from a list of 62319 Experts worldwide ranked by ideXlab platform

Sameer Maskey - One of the best experts on this subject based on the ideXlab platform.

  • speech segmentation and spoken Document Processing
    IEEE Signal Processing Magazine, 2008
    Co-Authors: Mari Ostendorf, Benoit Favre, Ralph Grishman, Dilek Hakkanitur, Mary P Harper, Dustin Hillard, Julia Hirschberg, Jeremy G Kahn, Yang Liu, Sameer Maskey
    Abstract:

    Progress in both speech and language Processing has spurred efforts to support applications that rely on spoken rather than written language input. A key challenge in moving from text-based Documents to such spoken Documents is that spoken language lacks explicit punctuation and formatting, which can be crucial for good performance. This article describes different levels of speech segmentation, approaches to automatically recovering segment boundary locations, and experimental results demonstrating impact on several language Processing tasks. The results also show a need for optimizing segmentation for the end task rather than independently.

Vincent Christlein - One of the best experts on this subject based on the ideXlab platform.

  • HDPA: historical Document Processing and analysis framework
    Evolving Systems, 2020
    Co-Authors: Ladislav Lenc, Jiří Martínek, Pavel Král, Anguelos Nicolao, Vincent Christlein
    Abstract:

    Nowadays, the accessibility of digitized historical Documents is extremely important to facilitate fast and efficient retrieval of historical information and knowledge extraction from such data. To provide such functionality, it is necessary to convert Document images into plain text using optical character recognition (OCR). Many OCR related methods and tools have been proposed, however, they are often too complicated for a standard user, some important parts are missing or they are not available in free versions. Therefore, this paper describes a complex and flexible web framework for historical Document manipulation and analysis with the main focus on OCR. The framework contains eight modules to facilitate three main tasks: image pre-Processing and segmentation, creation of data for OCR model training and the OCR itself. This framework is freely available for non commercial purposes. We have experimentally evaluated this framework on real data and we have shown that this system is efficient and can save human labour in the process of annotated data preparation. Moreover, we have reached state-of-the-art OCR results.

Mari Ostendorf - One of the best experts on this subject based on the ideXlab platform.

  • speech segmentation and spoken Document Processing
    IEEE Signal Processing Magazine, 2008
    Co-Authors: Mari Ostendorf, Benoit Favre, Ralph Grishman, Dilek Hakkanitur, Mary P Harper, Dustin Hillard, Julia Hirschberg, Jeremy G Kahn, Yang Liu, Sameer Maskey
    Abstract:

    Progress in both speech and language Processing has spurred efforts to support applications that rely on spoken rather than written language input. A key challenge in moving from text-based Documents to such spoken Documents is that spoken language lacks explicit punctuation and formatting, which can be crucial for good performance. This article describes different levels of speech segmentation, approaches to automatically recovering segment boundary locations, and experimental results demonstrating impact on several language Processing tasks. The results also show a need for optimizing segmentation for the end task rather than independently.

Benoit Favre - One of the best experts on this subject based on the ideXlab platform.

  • speech segmentation and spoken Document Processing
    IEEE Signal Processing Magazine, 2008
    Co-Authors: Mari Ostendorf, Benoit Favre, Ralph Grishman, Dilek Hakkanitur, Mary P Harper, Dustin Hillard, Julia Hirschberg, Jeremy G Kahn, Yang Liu, Sameer Maskey
    Abstract:

    Progress in both speech and language Processing has spurred efforts to support applications that rely on spoken rather than written language input. A key challenge in moving from text-based Documents to such spoken Documents is that spoken language lacks explicit punctuation and formatting, which can be crucial for good performance. This article describes different levels of speech segmentation, approaches to automatically recovering segment boundary locations, and experimental results demonstrating impact on several language Processing tasks. The results also show a need for optimizing segmentation for the end task rather than independently.

Mary P Harper - One of the best experts on this subject based on the ideXlab platform.

  • speech segmentation and spoken Document Processing
    IEEE Signal Processing Magazine, 2008
    Co-Authors: Mari Ostendorf, Benoit Favre, Ralph Grishman, Dilek Hakkanitur, Mary P Harper, Dustin Hillard, Julia Hirschberg, Jeremy G Kahn, Yang Liu, Sameer Maskey
    Abstract:

    Progress in both speech and language Processing has spurred efforts to support applications that rely on spoken rather than written language input. A key challenge in moving from text-based Documents to such spoken Documents is that spoken language lacks explicit punctuation and formatting, which can be crucial for good performance. This article describes different levels of speech segmentation, approaches to automatically recovering segment boundary locations, and experimental results demonstrating impact on several language Processing tasks. The results also show a need for optimizing segmentation for the end task rather than independently.