The Experts below are selected from a list of 798 Experts worldwide ranked by ideXlab platform

S Marcus - One of the best experts on this subject based on the ideXlab platform.

  • ICSLP - VPQ : A spoken language interface to large scale Directory Information
    1998
    Co-Authors: B Buntschuh, Candace A Kamm, G Di Fabbrizio, Alicia Abella, Mehryar Mohri, Shrikanth S Narayanan, Ilija Zeljkovic, R D Sharp, Jhg Wright, S Marcus
    Abstract:

    This paper describes VPQ (Voice Post Query), a dialog system that provides spoken access to the Information in the AT&T corporate personnel database (>120,000 entries). An explicit design goal is to have the user’s initial interaction with the system be rather unconstrained and to rely on tighter, prompt constrained, dialog only when absolutely necessary. The purpose of VPQ is both a) to explore and exploit the capabilities of “state of the art” speech recognition systems for this highperplexity task, and b) to develop the natural language understanding and dialog control components necessary for effective and efficient user interactions. The VPQ task spans a wide range of possible dialog scenarios. They range from simple “one-shot” to complex multi-turn interactions. The former correspond to interactions where the initial utterance is unambiguous and the system’s response appropriately terminates the interaction either by providing the desired Information or completing a call to the requested person. The more complex interactions occur primarily whenever ambiguities or errors require resolution. Current speech recognition accuracy of 80% is adequate to pursue such an ambitious task. This paper highlights the inherent challenges in such a task, the major components of the system, the rationale for their design, and how they perform. The VPQ project targets a variety of access devices, including telephony, desktop and handheld devices offering multi-modal user interfaces. In this paper we focus on describing the telephony interface.

  • vpq a spoken language interface to large scale Directory Information
    Conference of the International Speech Communication Association, 1998
    Co-Authors: B Buntschuh, Candace A Kamm, G Di Fabbrizio, Alicia Abella, Mehryar Mohri, Shrikanth S Narayanan, Ilija Zeljkovic, R D Sharp, Jhg Wright, S Marcus
    Abstract:

    This paper describes VPQ (Voice Post Query), a dialog system that provides spoken access to the Information in the AT&T corporate personnel database (>120,000 entries). An explicit design goal is to have the user’s initial interaction with the system be rather unconstrained and to rely on tighter, prompt constrained, dialog only when absolutely necessary. The purpose of VPQ is both a) to explore and exploit the capabilities of “state of the art” speech recognition systems for this highperplexity task, and b) to develop the natural language understanding and dialog control components necessary for effective and efficient user interactions. The VPQ task spans a wide range of possible dialog scenarios. They range from simple “one-shot” to complex multi-turn interactions. The former correspond to interactions where the initial utterance is unambiguous and the system’s response appropriately terminates the interaction either by providing the desired Information or completing a call to the requested person. The more complex interactions occur primarily whenever ambiguities or errors require resolution. Current speech recognition accuracy of 80% is adequate to pursue such an ambitious task. This paper highlights the inherent challenges in such a task, the major components of the system, the rationale for their design, and how they perform. The VPQ project targets a variety of access devices, including telephony, desktop and handheld devices offering multi-modal user interfaces. In this paper we focus on describing the telephony interface.

A. Kellner - One of the best experts on this subject based on the ideXlab platform.

  • padis an automatic telephone switchboard and Directory Information system
    Speech Communication, 1997
    Co-Authors: A. Kellner, B. Rueber, Frank Seide, Bachhiep Tran
    Abstract:

    PADIS, le systeme de standard telephonique automatique et d'Information annuaire de Philips offre une interface utilisateur en langage naturel pour acceder a une base de donnees telephoniques. En utilisant les technologies de reconnaissance de la parole et de comprehension de langage, le systeme permet d'obtenir les numeros de telephone, les numeros de fax, les adresses electroniques, les numeros de pieces ainsi que l'etablissement direct d'appel vers le numero desire. Dans cet article, nous presentons le cadre probabiliste sous-jacent, l'architecture du systeme, et les modules individuels de reconnaissance de parole, de comprehension du langage, de controle du dialogue, et de sortie vocale. De plus, nous rapportons des resultats sur les performances et le comportement des usagers obtenus a partir d'un test terrain realise dans notre laboratoire de recherche avec une base de donnees de 600 entrees. Nous derivons une nouvelle regle de decision basee sur le critere de maximum a posteriori qui incorpore des connaissances sur la base de donnees et sur l'historique du dialogue comme des contraintes pour la reconnaissance de la parole et la comprehension du langage. Ceci a permis d'ameliorer la comprehension de la parole de 19% (en termes de taux d'erreur), et de reduire de 38% les erreurs de substitution des attributs (par exemple reconnaissance d'un nom errone). La regle de decision est implantee dans une approche multiniveaux correspondant a une combinaison d'une reconnaissance de parole au niveau de l'etat de l'art, d'une recherche grammaticale partielle dans une grammaire a attributs hors contexte et stochastique, et d'un algorithme de recherche des N-meilleures solutions, qui est egalement decrit dans cet article. Le systeme conduit un dialogue d'initiative mixte flexible au lieu d'utiliser un schema rigide de remplissage de formulaires. Il incorpore des connaissances sur la base de donnees afin d'optimiser le deroulement du dialogue.

  • EUROSPEECH - Towards an automated Directory Information system.
    1997
    Co-Authors: Frank Seide, A. Kellner
    Abstract:

    This paper describes a design and feasibility study for a large-scale automatic Directory Information system with a scalable architecture. The current demonstrator, called PADIS-XL, operates in realtime and handles a database of a medium-size German city with 130,000 listings. The system uses a new technique of taking a combined decision on the joint probability over multiple dialogue turns, and a dialogue strategy that strives to restrict the search space more and more with every dialogue turn. During the course of the dialogue, the last name of the desired subscriber must be spelled out. The spelling recognizer permits continuous spelling and uses a context-free grammar to parse common spelling expressions. This paper describes the system architecture, our maximum a-posteriori (MAP) decision rule, the spelling grammar, and the dialogue strategy. We give results on the SPEECHDAT and SIETILL databases on recognition of first names by spelling and on jointly deciding on the spelled and the spoken name. In a 35,000-names setup, the joint decision reduced name-recognition errors by 31%.

  • towards an automated Directory Information system
    Conference of the International Speech Communication Association, 1997
    Co-Authors: Frank Seide, A. Kellner
    Abstract:

    This paper describes a design and feasibility study for a large-scale automatic Directory Information system with a scalable architecture. The current demonstrator, called PADIS-XL, operates in realtime and handles a database of a medium-size German city with 130,000 listings. The system uses a new technique of taking a combined decision on the joint probability over multiple dialogue turns, and a dialogue strategy that strives to restrict the search space more and more with every dialogue turn. During the course of the dialogue, the last name of the desired subscriber must be spelled out. The spelling recognizer permits continuous spelling and uses a context-free grammar to parse common spelling expressions. This paper describes the system architecture, our maximum a-posteriori (MAP) decision rule, the spelling grammar, and the dialogue strategy. We give results on the SPEECHDAT and SIETILL databases on recognition of first names by spelling and on jointly deciding on the spelled and the spoken name. In a 35,000-names setup, the joint decision reduced name-recognition errors by 31%.

  • A VOICE-CONTROLLED AUTOMATIC TELEPH SWITCHBOARD AND Directory Information S
    1996
    Co-Authors: A. Kellner, B. Rueber, F. Seide
    Abstract:

    In this paper, we present the Philips automatic telephone switchboard and Directory Information system PADIS'. PADIS understands natural-language requests in fluently spoken German. The system offers telephone, fax, and room numbers, email addresses, private phone numbers, and direct call completion. A setup with a 500-entry database is currently in a field test in our research laboratory and has shown a success rate of 90%. This paper describes the system architecture and its components, and presents experiences as well as results from the field test.

  • A voice-controlled automatic telephone switchboard and Directory Information system
    Proceedings of IVTTA '96. Workshop on Interactive Voice Technology for Telecommunications Applications, 1996
    Co-Authors: A. Kellner, B. Rueber, F. Seide
    Abstract:

    In this paper, we present the Philips automatic telephone switchboard and Directory Information system PADIS. PADIS understands natural-language requests in fluently spoken German. The system offers telephone, fax, and room numbers, email addresses, private phone numbers, and direct call completion. A setup with a 500-entry database is currently in a field test in our research laboratory and has shown a success rate of 90%. This paper describes the system architecture and its components, and presents experiences as well as results from the field test.

F. Seide - One of the best experts on this subject based on the ideXlab platform.

  • With a little help from the database-developing voice-controlled Directory Information systems
    1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings, 1997
    Co-Authors: A. Kellner, F. Seide, B. Rueber
    Abstract:

    Automated Directory Information is amongst the most challenging applications of automatic speech recognition. We present some basic techniques that try to overcome the deficiencies of the speech recognizer by incorporating as much additional knowledge as possible, such as the telephone Directory. We derive a maximum a-posteriori decision rule which explicitly uses the telephone Directory knowledge as well as the dialogue history to improve speech understanding accuracy. The rule allows us to take a combined decision on the joint probability over multiple dialogue turns, which yields good results in combination with spelling. Our spelling architecture permits continuous spelling of names and uses a context-free grammar to parse common spelling expressions. We review two different real time prototypes, on which we evaluated our decision rule. One (PADIS) operates on a small database and one (PADIS-XL) on a database with 130000 entries.

  • A VOICE-CONTROLLED AUTOMATIC TELEPH SWITCHBOARD AND Directory Information S
    1996
    Co-Authors: A. Kellner, B. Rueber, F. Seide
    Abstract:

    In this paper, we present the Philips automatic telephone switchboard and Directory Information system PADIS'. PADIS understands natural-language requests in fluently spoken German. The system offers telephone, fax, and room numbers, email addresses, private phone numbers, and direct call completion. A setup with a 500-entry database is currently in a field test in our research laboratory and has shown a success rate of 90%. This paper describes the system architecture and its components, and presents experiences as well as results from the field test.

  • A voice-controlled automatic telephone switchboard and Directory Information system
    Proceedings of IVTTA '96. Workshop on Interactive Voice Technology for Telecommunications Applications, 1996
    Co-Authors: A. Kellner, B. Rueber, F. Seide
    Abstract:

    In this paper, we present the Philips automatic telephone switchboard and Directory Information system PADIS. PADIS understands natural-language requests in fluently spoken German. The system offers telephone, fax, and room numbers, email addresses, private phone numbers, and direct call completion. A setup with a 500-entry database is currently in a field test in our research laboratory and has shown a success rate of 90%. This paper describes the system architecture and its components, and presents experiences as well as results from the field test.

B Buntschuh - One of the best experts on this subject based on the ideXlab platform.

  • ICSLP - VPQ : A spoken language interface to large scale Directory Information
    1998
    Co-Authors: B Buntschuh, Candace A Kamm, G Di Fabbrizio, Alicia Abella, Mehryar Mohri, Shrikanth S Narayanan, Ilija Zeljkovic, R D Sharp, Jhg Wright, S Marcus
    Abstract:

    This paper describes VPQ (Voice Post Query), a dialog system that provides spoken access to the Information in the AT&T corporate personnel database (>120,000 entries). An explicit design goal is to have the user’s initial interaction with the system be rather unconstrained and to rely on tighter, prompt constrained, dialog only when absolutely necessary. The purpose of VPQ is both a) to explore and exploit the capabilities of “state of the art” speech recognition systems for this highperplexity task, and b) to develop the natural language understanding and dialog control components necessary for effective and efficient user interactions. The VPQ task spans a wide range of possible dialog scenarios. They range from simple “one-shot” to complex multi-turn interactions. The former correspond to interactions where the initial utterance is unambiguous and the system’s response appropriately terminates the interaction either by providing the desired Information or completing a call to the requested person. The more complex interactions occur primarily whenever ambiguities or errors require resolution. Current speech recognition accuracy of 80% is adequate to pursue such an ambitious task. This paper highlights the inherent challenges in such a task, the major components of the system, the rationale for their design, and how they perform. The VPQ project targets a variety of access devices, including telephony, desktop and handheld devices offering multi-modal user interfaces. In this paper we focus on describing the telephony interface.

  • vpq a spoken language interface to large scale Directory Information
    Conference of the International Speech Communication Association, 1998
    Co-Authors: B Buntschuh, Candace A Kamm, G Di Fabbrizio, Alicia Abella, Mehryar Mohri, Shrikanth S Narayanan, Ilija Zeljkovic, R D Sharp, Jhg Wright, S Marcus
    Abstract:

    This paper describes VPQ (Voice Post Query), a dialog system that provides spoken access to the Information in the AT&T corporate personnel database (>120,000 entries). An explicit design goal is to have the user’s initial interaction with the system be rather unconstrained and to rely on tighter, prompt constrained, dialog only when absolutely necessary. The purpose of VPQ is both a) to explore and exploit the capabilities of “state of the art” speech recognition systems for this highperplexity task, and b) to develop the natural language understanding and dialog control components necessary for effective and efficient user interactions. The VPQ task spans a wide range of possible dialog scenarios. They range from simple “one-shot” to complex multi-turn interactions. The former correspond to interactions where the initial utterance is unambiguous and the system’s response appropriately terminates the interaction either by providing the desired Information or completing a call to the requested person. The more complex interactions occur primarily whenever ambiguities or errors require resolution. Current speech recognition accuracy of 80% is adequate to pursue such an ambitious task. This paper highlights the inherent challenges in such a task, the major components of the system, the rationale for their design, and how they perform. The VPQ project targets a variety of access devices, including telephony, desktop and handheld devices offering multi-modal user interfaces. In this paper we focus on describing the telephony interface.

B. Rueber - One of the best experts on this subject based on the ideXlab platform.

  • With a little help from the database-developing voice-controlled Directory Information systems
    1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings, 1997
    Co-Authors: A. Kellner, F. Seide, B. Rueber
    Abstract:

    Automated Directory Information is amongst the most challenging applications of automatic speech recognition. We present some basic techniques that try to overcome the deficiencies of the speech recognizer by incorporating as much additional knowledge as possible, such as the telephone Directory. We derive a maximum a-posteriori decision rule which explicitly uses the telephone Directory knowledge as well as the dialogue history to improve speech understanding accuracy. The rule allows us to take a combined decision on the joint probability over multiple dialogue turns, which yields good results in combination with spelling. Our spelling architecture permits continuous spelling of names and uses a context-free grammar to parse common spelling expressions. We review two different real time prototypes, on which we evaluated our decision rule. One (PADIS) operates on a small database and one (PADIS-XL) on a database with 130000 entries.