The Experts below are selected from a list of 10281 Experts worldwide ranked by ideXlab platform

Moira Regelson - One of the best experts on this subject based on the ideXlab platform.

  • the Linguistic Structure of english web search queries
    Empirical Methods in Natural Language Processing, 2008
    Co-Authors: Cory Barr, Rosie Jones, Moira Regelson
    Abstract:

    Web-search queries are known to be short, but little else is known about their Structure. In this paper we investigate the applicability of part-of-speech tagging to typical English-language web search-engine queries and the potential value of these tags for improving search results. We begin by identifying a set of part-of-speech tags suitable for search queries and quantifying their occurrence. We find that proper-nouns constitute 40% of query terms, and proper nouns and nouns together constitute over 70% of query terms. We also show that the majority of queries are noun-phrases, not unStructured collections of terms. We then use a set of queries manually labeled with these tags to train a Brill tagger and evaluate its performance. In addition, we investigate classification of search queries into grammatical classes based on the syntax of part-of-speech tag sequences. We also conduct preliminary investigative experiments into the practical applicability of leveraging query-trained part-of-speech taggers for information-retrieval tasks. In particular, we show that part-of-speech information can be a significant feature in machine-learned search-result relevance. These experiments also include the potential use of the tagger in selecting words for omission or substitution in query reformulation, actions which can improve recall. We conclude that training a part-of-speech tagger on labeled corpora of queries significantly outperforms taggers based on traditional corpora, and leveraging the unique Linguistic Structure of web-search queries can improve search experience.

  • EMNLP - The Linguistic Structure of English Web-Search Queries
    Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08, 2008
    Co-Authors: Cory Barr, Rosie Jones, Moira Regelson
    Abstract:

    Web-search queries are known to be short, but little else is known about their Structure. In this paper we investigate the applicability of part-of-speech tagging to typical English-language web search-engine queries and the potential value of these tags for improving search results. We begin by identifying a set of part-of-speech tags suitable for search queries and quantifying their occurrence. We find that proper-nouns constitute 40% of query terms, and proper nouns and nouns together constitute over 70% of query terms. We also show that the majority of queries are noun-phrases, not unStructured collections of terms. We then use a set of queries manually labeled with these tags to train a Brill tagger and evaluate its performance. In addition, we investigate classification of search queries into grammatical classes based on the syntax of part-of-speech tag sequences. We also conduct preliminary investigative experiments into the practical applicability of leveraging query-trained part-of-speech taggers for information-retrieval tasks. In particular, we show that part-of-speech information can be a significant feature in machine-learned search-result relevance. These experiments also include the potential use of the tagger in selecting words for omission or substitution in query reformulation, actions which can improve recall. We conclude that training a part-of-speech tagger on labeled corpora of queries significantly outperforms taggers based on traditional corpora, and leveraging the unique Linguistic Structure of web-search queries can improve search experience.

Jonathan Brennan - One of the best experts on this subject based on the ideXlab platform.

  • phase synchronization varies systematically with Linguistic Structure composition
    Philosophical Transactions of the Royal Society B, 2020
    Co-Authors: Jonathan Brennan, Andrea E. Martin
    Abstract:

    Computation in neuronal assemblies is putatively reflected in the excitatory and inhibitory cycles of activation distributed throughout the brain. In speech and language processing, coordination of...

  • modeling fmri time courses with Linguistic Structure at various grain sizes
    North American Chapter of the Association for Computational Linguistics, 2015
    Co-Authors: John Hale, David Lutz, Jonathan Brennan
    Abstract:

    Neuroimaging while participants listen to audiobooks provides a rich data source for theories of incremental parsing. We compare nested regression models of these data. These mixed-effects models incorporate Linguistic predictors at various grain sizes ranging from part-of-speech bigrams, through surprisal on context-free treebank grammars, to incremental node counts in trees that are derived by Minimalist Grammars. The fine-grained Structures make an independent contribution over and above coarser predictors. However, this result only obtains with time courses from anterior temporal lobe (aTL). In analogous time courses from inferior frontal gyrus, only n-grams improve upon a non-syntactic baseline. These results support the idea that aTL does combinatoric processing during naturalistic story comprehension, processing that bears a systematic relationship to Linguistic Structure.

  • CMCL@NAACL-HLT - Modeling fMRI time courses with Linguistic Structure at various grain sizes
    Proceedings of the 6th Workshop on Cognitive Modeling and Computational Linguistics, 2015
    Co-Authors: John Hale, David Lutz, Jonathan Brennan
    Abstract:

    Neuroimaging while participants listen to audiobooks provides a rich data source for theories of incremental parsing. We compare nested regression models of these data. These mixed-effects models incorporate Linguistic predictors at various grain sizes ranging from part-of-speech bigrams, through surprisal on context-free treebank grammars, to incremental node counts in trees that are derived by Minimalist Grammars. The fine-grained Structures make an independent contribution over and above coarser predictors. However, this result only obtains with time courses from anterior temporal lobe (aTL). In analogous time courses from inferior frontal gyrus, only n-grams improve upon a non-syntactic baseline. These results support the idea that aTL does combinatoric processing during naturalistic story comprehension, processing that bears a systematic relationship to Linguistic Structure.

Fey Parrill - One of the best experts on this subject based on the ideXlab platform.

  • viewpoint in speech gesture integration Linguistic Structure discourse Structure and event Structure
    Language and Cognitive Processes, 2010
    Co-Authors: Fey Parrill
    Abstract:

    We examine a corpus of narrative data to determine which types of events evoke character viewpoint gestures, and which evoke observer viewpoint gestures. We consider early claims made by McNeill (1992) that character viewpoint tends to occur with transitive utterances and utterances that are causally central to the narrative. We argue that the Structure of the event itself must be taken into account: there are some events that cannot plausibly evoke both types of gesture. We show that Linguistic Structure (transitivity), event Structure (visuo-spatial and motoric properties), and discourse Structure all play a role. We apply these findings to a recent model of embodied language production, the Gestures as Simulated Action framework.

  • Viewpoint in speech–gesture integration: Linguistic Structure, discourse Structure, and event Structure
    Language and Cognitive Processes, 2010
    Co-Authors: Fey Parrill
    Abstract:

    We examine a corpus of narrative data to determine which types of events evoke character viewpoint gestures, and which evoke observer viewpoint gestures. We consider early claims made by McNeill (1992) that character viewpoint tends to occur with transitive utterances and utterances that are causally central to the narrative. We argue that the Structure of the event itself must be taken into account: there are some events that cannot plausibly evoke both types of gesture. We show that Linguistic Structure (transitivity), event Structure (visuo-spatial and motoric properties), and discourse Structure all play a role. We apply these findings to a recent model of embodied language production, the Gestures as Simulated Action framework.

Cory Barr - One of the best experts on this subject based on the ideXlab platform.

  • the Linguistic Structure of english web search queries
    Empirical Methods in Natural Language Processing, 2008
    Co-Authors: Cory Barr, Rosie Jones, Moira Regelson
    Abstract:

    Web-search queries are known to be short, but little else is known about their Structure. In this paper we investigate the applicability of part-of-speech tagging to typical English-language web search-engine queries and the potential value of these tags for improving search results. We begin by identifying a set of part-of-speech tags suitable for search queries and quantifying their occurrence. We find that proper-nouns constitute 40% of query terms, and proper nouns and nouns together constitute over 70% of query terms. We also show that the majority of queries are noun-phrases, not unStructured collections of terms. We then use a set of queries manually labeled with these tags to train a Brill tagger and evaluate its performance. In addition, we investigate classification of search queries into grammatical classes based on the syntax of part-of-speech tag sequences. We also conduct preliminary investigative experiments into the practical applicability of leveraging query-trained part-of-speech taggers for information-retrieval tasks. In particular, we show that part-of-speech information can be a significant feature in machine-learned search-result relevance. These experiments also include the potential use of the tagger in selecting words for omission or substitution in query reformulation, actions which can improve recall. We conclude that training a part-of-speech tagger on labeled corpora of queries significantly outperforms taggers based on traditional corpora, and leveraging the unique Linguistic Structure of web-search queries can improve search experience.

  • EMNLP - The Linguistic Structure of English Web-Search Queries
    Proceedings of the Conference on Empirical Methods in Natural Language Processing - EMNLP '08, 2008
    Co-Authors: Cory Barr, Rosie Jones, Moira Regelson
    Abstract:

    Web-search queries are known to be short, but little else is known about their Structure. In this paper we investigate the applicability of part-of-speech tagging to typical English-language web search-engine queries and the potential value of these tags for improving search results. We begin by identifying a set of part-of-speech tags suitable for search queries and quantifying their occurrence. We find that proper-nouns constitute 40% of query terms, and proper nouns and nouns together constitute over 70% of query terms. We also show that the majority of queries are noun-phrases, not unStructured collections of terms. We then use a set of queries manually labeled with these tags to train a Brill tagger and evaluate its performance. In addition, we investigate classification of search queries into grammatical classes based on the syntax of part-of-speech tag sequences. We also conduct preliminary investigative experiments into the practical applicability of leveraging query-trained part-of-speech taggers for information-retrieval tasks. In particular, we show that part-of-speech information can be a significant feature in machine-learned search-result relevance. These experiments also include the potential use of the tagger in selecting words for omission or substitution in query reformulation, actions which can improve recall. We conclude that training a part-of-speech tagger on labeled corpora of queries significantly outperforms taggers based on traditional corpora, and leveraging the unique Linguistic Structure of web-search queries can improve search experience.

Chun-an Chan - One of the best experts on this subject based on the ideXlab platform.

  • Unsupervised discovery of Linguistic Structure including two-level acoustic patterns using three cascaded stages of iterative optimization
    2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013
    Co-Authors: Cheng-tao Chung, Chun-an Chan
    Abstract:

    Techniques for unsupervised discovery of acoustic patterns are getting increasingly attractive, because huge quantities of speech data are becoming available but manual annotations remain hard to acquire. In this paper, we propose an approach for unsupervised discovery of Linguistic Structure for the target spoken language given raw speech data. This Linguistic Structure includes two-level (subword-like and word-like) acoustic patterns, the lexicon of word-like patterns in terms of subword-like patterns and the N-gram language model based on word-like patterns. All patterns, models, and parameters can be automatically learned from the unlabelled speech corpus. This is achieved by an initialization step followed by three cascaded stages for acoustic, Linguistic, and lexical iterative optimization. The lexicon of word-like patterns defines allowed consecutive sequence of HMMs for subword-like patterns. In each iteration, model training and decoding produces updated labels from which the lexicon and HMMs can be further updated. In this way, model parameters and decoded labels are respectively optimized in each iteration, and the knowledge about the Linguistic Structure is learned gradually layer after layer. The proposed approach was tested in preliminary experiments on a corpus of Mandarin broadcast news, including a task of spoken term detection with performance compared to a parallel test using models trained in a supervised way. Results show that the proposed system not only yields reasonable performance on its own, but is also complimentary to existing large vocabulary ASR systems.

  • ICASSP - Unsupervised discovery of Linguistic Structure including two-level acoustic patterns using three cascaded stages of iterative optimization
    2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013
    Co-Authors: Cheng-tao Chung, Chun-an Chan
    Abstract:

    Techniques for unsupervised discovery of acoustic patterns are getting increasingly attractive, because huge quantities of speech data are becoming available but manual annotations remain hard to acquire. In this paper, we propose an approach for unsupervised discovery of Linguistic Structure for the target spoken language given raw speech data. This Linguistic Structure includes two-level (subword-like and word-like) acoustic patterns, the lexicon of word-like patterns in terms of subword-like patterns and the N-gram language model based on word-like patterns. All patterns, models, and parameters can be automatically learned from the unlabelled speech corpus. This is achieved by an initialization step followed by three cascaded stages for acoustic, Linguistic, and lexical iterative optimization. The lexicon of word-like patterns defines allowed consecutive sequence of HMMs for subword-like patterns. In each iteration, model training and decoding produces updated labels from which the lexicon and HMMs can be further updated. In this way, model parameters and decoded labels are respectively optimized in each iteration, and the knowledge about the Linguistic Structure is learned gradually layer after layer. The proposed approach was tested in preliminary experiments on a corpus of Mandarin broadcast news, including a task of spoken term detection with performance compared to a parallel test using models trained in a supervised way. Results show that the proposed system not only yields reasonable performance on its own, but is also complimentary to existing large vocabulary ASR systems.