The Experts below are selected from a list of 4593 Experts worldwide ranked by ideXlab platform
James R Curran - One of the best experts on this subject based on the ideXlab platform.
-
improving combinatory Categorial Grammar parse reranking with dependency Grammar features
International Conference on Computational Linguistics, 2012Co-Authors: Sunghwan Mac Kim, Mark Johnson, James R CurranAbstract:This paper presents a novel method of improving Combinatory Categorial Grammar (CCG) parsing using features generated from Dependency Grammar (DG) parses and combined using reranking. Different Grammar formalisms have different strengths and different parsing models have consequently divergent views of the data. More specifically, dependency parsers are sensitive to linguistic generalisations that differ from the generalisations that the CCG parser is sensitive to, and which the reranker exploits to identify the parse most likely to be correct. We propose DG-derived reranking features, which are obtained by comparing dependencies from the CCG parser with DG dependencies, and demonstrate how they improve the performance of a CCG parser and reranker in a variety of settings. We record a final labeled F-score of 87.93% on section 23 of CCGbank, 0.5% and 0.35% improvements over the base parser (87.43%) and reranker (87.58%), respectively.
-
the challenges of parsing chinese with combinatory Categorial Grammar
North American Chapter of the Association for Computational Linguistics, 2012Co-Authors: Daniel Tse, James R CurranAbstract:We apply Combinatory Categorial Grammar to wide-coverage parsing in Chinese with the new Chinese CCGbank, bringing a formalism capable of transparently recovering non-local dependencies to a language in which they are particularly frequent. We train two state-of-the-art English CCG parsers: the parser of Petrov and Klein (P&K), and the Clark and Curran (C&C) parser, uncovering a surprising performance gap between them not observed in English --- 72.73 (P&K) and 67.09 (C&C) F-score on PCTB 6. We explore the challenges of Chinese CCG parsing through three novel ideas: developing corpus variants rather than treating the corpus as fixed; controlling noun/verb and other POS ambiguities; and quantifying the impact of constructions like pro-drop.
-
efficient combinatory Categorial Grammar parsing
Proceedings of the Australasian Language Technology Workshop 2006, 2006Co-Authors: Bojan Djordjevic, James R CurranAbstract:Efficient wide-coverage parsing is integral to large-scale NLP applications. Unfortunately, parsers for linguistically motivated formalisms, e.g. HPSG and TAG, are often too inefficient for these applications. This paper describes two modifications to the standard CKY chart parsing algorithm used in the Clark and Curran (2006) Combinatory Categorial Grammar (CCG) parser. The first modification extends the tight integration of the supertagger and parser, so that individual supertags can be added to the chart, which is then repaired rather than rebuilt. The second modification adds constraints to the chart that restrict which constituents can combine. Parsing speed is improved by 30–35% without a significant accuracy penalty and a small increase in coverage when both of these modifications are used.
Yusuke Kubota - One of the best experts on this subject based on the ideXlab platform.
-
development of a general purpose Categorial Grammar treebank
Language Resources and Evaluation, 2020Co-Authors: Yusuke Kubota, Koji Mineshima, Noritsugu Hayashi, Shinya OkanoAbstract:This paper introduces ABC Treebank, a general-purpose Categorial Grammar (CG) treebank for Japanese. It is ‘general-purpose’ in the sense that it is not tailored to a specific variant of CG, but rather aims to offer a theory-neutral linguistic resource (as much as possible) which can be converted to different versions of CG (specifically, CCG and Type-Logical Grammar) relatively easily. In terms of linguistic analysis, it improves over the existing Japanese CG treebank (Japanese CCGBank) on the treatment of certain linguistic phenomena (passives, causatives, and control/raising predicates) for which the lexical specification of the syntactic information reflecting local dependencies turns out to be crucial. In this paper, we describe the underlying ‘theory’ dubbed ABC Grammar that is taken as a basis for our treebank, outline the general construction of the corpus, and report on some preliminary results applying the treebank in a semantic parsing system for generating logical representations of sentences.
-
Pseudogapping as Pseudo-VP ellipsis
2020Co-Authors: Yusuke Kubota, Robert LevineAbstract:Abstract. In this paper, we propose an analysis of pseudogapping in Hybrid Type-Logical Categorial Grammar (Hybrid TLCG
-
Pseudogapping as pseudo-VP ellipsis
2020Co-Authors: Yusuke Kubota, Robert LevineAbstract:Abstract In this paper, we propose an analysis of pseudogapping in Hybrid Type-Logical Categorial Grammar (Hybrid TLCG; Kubota 2010; Kubota and Levine 2012). Pseudogapping poses a particularly challenging problem for previous analyses in both the transformational and nontransformational literature. We argue that the flexible syntaxsemantics interface of Hybrid TLCG enables an analysis of pseudogapping that synthesizes the key insights of both transformational and nontransformational approaches, while at the same time overcoming the major difficulties of each type of approach. Keywords: pseudogapping, VP ellipsis, anaphora, syntactic identity, Hybrid Type-Logical Categorial Grammar 0 We are indebted to the following people for comments and discussions: Ai Kubota, Scott Martin, Philip Miller, Jordan Needle, Carl Pollard, and Daniel Puthawala. We would also like to thank the two LI reviewers for their insightful comments. We have presented parts of the present work at various venues, including LACL 2014, the syntax semantics group Synners at OSU, and the Semantics Workshop in Tokai (Nagoya). We would like to thank the audiences at these venues for their feedback
-
The syntax-semantics interface of ‘respective’ predication: a unified analysis in Hybrid Type-Logical Categorial Grammar
Natural Language & Linguistic Theory, 2016Co-Authors: Yusuke Kubota, Robert LevineAbstract:This paper proposes a unified analysis of the ‘respective’ readings of plural and conjoined expressions, the internal readings of symmetrical predicates such as same and different , and the summative readings of expressions such as a total of $10000 . These expressions pose significant challenges to compositional semantics, and have been studied extensively in the literature. However, almost all previous studies focus exclusively on one of these phenomena, and the close parallels and interactions that they exhibit have been mostly overlooked to date. We point out two key properties common to these phenomena: (i) they target all types of coordination, including nonconstituent coordination such as Right-Node Raising and Dependent Cluster Coordination; (ii) the three phenomena all exhibit multiple dependency, both by themselves and with respect to each other. These two parallels suggest that one and the same mechanism is at the core of their semantics. Building on this intuition, we propose a unified analysis of these phenomena, in which the meanings of expressions involving coordination are formally modelled as multisets, that is, sets that allow for duplicate occurrences of identical elements. The analysis is couched in Hybrid Type-Logical Categorial Grammar. The flexible syntax-semantics interface of this framework enables an analysis of ‘respective’ readings and related phenomena which, for the first time in the literature, yields a simple and principled solution for both the interactions with nonconstituent coordination and the multiple dependency noted above.
-
nonconstituent coordination in japanese as constituent coordination an analysis in hybrid type logical Categorial Grammar
Linguistic Inquiry, 2015Co-Authors: Yusuke KubotaAbstract:Nonconstituent coordination poses a particularly challenging problem for standard kinds of syntactic theories in which the notion of phrase structure (or constituency) is taken to be a primitive in some way or other. Previous approaches within such theories essentially equate nonconstituent coordination with coordination of full-fledged clauses at some level of grammatical representation. I present data from Japanese that pose problems for such approaches and argue for an alternative analysis in which the apparent nonconstituents are in fact surface constituents having full-fledged meanings, couched in a framework called Hybrid Type-Logical Categorial Grammar (Kubota 2010, Kubota and Levine 2012, Kubota 2014).
Julia Hockenmaier - One of the best experts on this subject based on the ideXlab platform.
-
induction of linguistic structure with combinatory Categorial Grammars
North American Chapter of the Association for Computational Linguistics, 2012Co-Authors: Yonatan Bisk, Julia HockenmaierAbstract:Our system consists of a simple, EM-based induction algorithm (Bisk and Hockenmaier, 2012), which induces a language-specific Combinatory Categorial Grammar (CCG) and lexicon based on a small number of linguistic principles, e.g. that verbs may be the roots of sentences and can take nouns as arguments.
-
priming effects in combinatory Categorial Grammar
Empirical Methods in Natural Language Processing, 2006Co-Authors: David Reitter, Julia Hockenmaier, Frank KellerAbstract:This paper presents a corpus-based account of structural priming in human sentence processing, focusing on the role that syntactic representations play in such an account. We estimate the strength of structural priming effects from a corpus of spontaneous spoken dialogue, annotated syntactically with Combinatory Categorial Grammar (CCG) derivations. This methodology allows us to test a range of predictions that CCG makes about priming. In particular, we present evidence for priming between lexical and syntactic categories encoding partially satisfied sub-categorization frames, and we show that priming effects exist both for incremental and normal-form CCG derivations.
-
identifying semantic roles using combinatory Categorial Grammar
Empirical Methods in Natural Language Processing, 2003Co-Authors: Daniel Gildea, Julia HockenmaierAbstract:We present a system for automatically identifying PropBank-style semantic roles based on the output of a statistical parser for Combinatory Categorial Grammar. This system performs at least as well as a system based on a traditional Treebank parser, and outperforms it on core argument roles.
-
generative models for statistical parsing with combinatory Categorial Grammar
Meeting of the Association for Computational Linguistics, 2002Co-Authors: Julia Hockenmaier, Mark SteedmanAbstract:This paper compares a number of generative probability models for a wide-coverage Combinatory Categorial Grammar (CCG) parser. These models are trained and tested on a corpus obtained by translating the Penn Treebank trees into CCG normal-form derivations. According to an evaluation of unlabeled word-word dependencies, our best model achieves a performance of 89.9%, comparable to the figures given by Collins (1999) for a linguistically less expressive Grammar. In contrast to Gildea (2001), we find a significant improvement from modeling word-word dependencies.
Jong Cheol Park - One of the best experts on this subject based on the ideXlab platform.
-
text parsing for sign language generation with combinatory Categorial Grammar
2011Co-Authors: Jinwoo Chung, Jong Cheol ParkAbstract:paper, we propose a method to convert a written sentence in spoken language into a suitable representation in sign language within the framework of Combinatory Categorial Grammar (CCG). The representation reflects the multi-channel nature of sign language performance, including manual and non-manual linguistic signals of multiple channels and information about their coordination. We show that most information needed to address linguistic phenomena in sign language such as word order, spatial references, classifier construction, and verb inflection can be encoded in the CCG sign lexicon. During the CCG derivation process, a semantic representation for sign language expressions is created so that the resulting output can be directly interpreted as a sequence of signs, each containing manual and non-manual components and representing their coordination and spatial relationship. The derivation process with the constructed lexicon is presented with several examples for Korean Sign Language. We discuss implications of our proposal and future directions.
-
interpretation of natural language queries for relational database access with combinatory Categorial Grammar
International Journal of Computer Processing of Languages, 2002Co-Authors: Hodong Lee, Jong Cheol ParkAbstract:In this paper, we describe a proposal to derive formal language queries from natural language queries with a combinatory Categorial Grammar (CCG). CCGs are well known to provide a means of deriving all the levels of information for natural language, i.e., syntax, semantics and discourse, at the same time. In our proposal, we utilize an extra level of representation for formal language queries for the aforementioned derivation. The syntactic coverage is shown with various natural language queries, including compound nouns, modification markers, various types of ellipses, numerical expressions, and subordinate and coordinate constructions. The general purpose CCG lexicon is semi-automatically augmented with the database fields and entries. We also discuss the performance of an implemented natural language query processing system.
-
using combinatory Categorial Grammar to extract biomedical information
IEEE Intelligent Systems, 2001Co-Authors: Jong Cheol ParkAbstract:Extracting information from biology databases manually can be an overwhelming task. GenBank, the US National Institutes of Health database containing all publicly available DNA sequences, has more than 14 billion bases in 13 million genetic-sequence records. Medline, a literature database available through PubMed, has over 11 million journal citations. In a May 2001 search request for "cytokine" (regulatory proteins in the immune system), PubMed returned 296556 articles. Given the quantity and complexity of biomedical literature, demands for computational tools to extract specific information are increasing. The author reviews biomedical information extraction methods and presents research done by KAIST's natural language processing group on a system that shows encouraging performance using combinatory Categorial Grammar as a natural language Grammar formalism.
-
informed parsing for coordination with combinatory Categorial Grammar
International Conference on Computational Linguistics, 2000Co-Authors: Jong Cheol Park, Hyung Joon ChoAbstract:Coordination in natural language hampers efficient parsing, especially due to the multiple and mostly unintended candidate conjuncts/disjuncts in a given sentence that shows structural ambiguity. The problem gets more serious in a combinatory Categorial Grammar framework, which is well known for its competent treatment of coordination, as the flexibility of syntactic analysis often strikes back as spurious ambiguity. We propose to address these ambiguities with predicate argument structures and semantic co-occurrence similarity information, and present encouraging results.
-
combinatory Categorial Grammar for the syntactic semantic and discourse analyses of coordinate constructions in korean
Journal of KIISE:Software and Applications, 2000Co-Authors: Hyung Joon Cho, Jong Cheol ParkAbstract:Coordinate constructions in natural language pose a number of difficulties to natural language processing units, due to the increased complexity of syntactic analysis, the syntactic ambiguity of the involved lexical items, and the apparent deletion of predicates in various places. In this paper, we address the syntactic characteristics of the coordinate constructions in Korean from the viewpoint of constructing a competence Grammar, and present a version of combinatory Categorial Grammar for the analysis of coordinate constructions in Korean. We also show how to utilize a unified lexicon in the proposed Grammar formalism in deriving the sentential semantics and associated information structures as well, in order to capture the discourse functions of coordinate constructions in Korean. The presented analysis conforms to the common wisdom that coordinate constructions are utilized in language not simply to reduce multiple sentences to a single sentence, but also to convey the information of contrast. Finally, we provide an analysis of sample corpora for the frequency of coordinate constructions in Korean and discuss some problematic cases.
Peter Schuller - One of the best experts on this subject based on the ideXlab platform.
-
flexible combinatory Categorial Grammar parsing using the cyk algorithm and answer set programming
International Conference on Logic Programming, 2013Co-Authors: Peter SchullerAbstract:Combinatory Categorial Grammar CCG is a Grammar formalism used for natural language parsing. CCG assigns structured lexical categories to words and uses a small set of combinatory rules to combine these categories in order to parse sentences. In this work we describe and implement a new approach to CCG parsing that relies on Answer Set Programming ASP -- a declarative programming paradigm.Different from previous work, we present an encoding that is inspired by the algorithm due to Cocke, Younger, and Kasami CYK. We also show encoding extensions for parse tree normalization and best-effort parsing and outline possible future extensions which are possible due to the usage of ASP as computational mechanism. We analyze performance of our approach on a part of the Brown corpus and discuss lessons learned during experiments with the ASP tools dlv, gringo, and clasp. The new approach is available in the open source CCG parsing toolkit AspCcgTk which uses the C&C supertagger as a preprocessor to achieve wide-coverage natural language parsing.
-
parsing combinatory Categorial Grammar via planning in answer set programming
Correct Reasoning, 2012Co-Authors: Yuliya Lierler, Peter SchullerAbstract:Combinatory Categorial Grammar (CCG) is a Grammar formalism used for natural language parsing. CCG assigns structured lexical categories to words and uses combinatory rules to combine these categories to parse a sentence. In this work we propose and implement a new approach to CCG parsing that relies on a prominent knowledge representation formalism, answer set programming (ASP) - a declarative programming paradigm. We formulate the task of CCG parsing as a planning problem and use an ASP computational tool to compute solutions that correspond to valid parses. Compared to other approaches, there is no need to implement a specific parsing algorithm using such a declarative method. Our approach aims at producing all semantically distinct parse trees for a given sentence. From this goal, normalization and efficiency issues arise, and we deal with them by combining and extending existing strategies.We have implemented a CCG parsing tool kit-AspCcgTk-that uses ASP as its main computational means. The C&C supertagger can be used as a preprocessor within AspCcgTk, which allows us to achieve wide-coverage natural language parsing.
-
parsing combinatory Categorial Grammar with answer set programming preliminary report
arXiv: Artificial Intelligence, 2011Co-Authors: Yuliya Lierler, Peter SchullerAbstract:Combinatory Categorial Grammar (CCG) is a Grammar formalism used for natural language parsing. CCG assigns structured lexical categories to words and uses a small set of combinatory rules to combine these categories to parse a sentence. In this work we propose and implement a new approach to CCG parsing that relies on a prominent knowledge representation formalism, answer set programming (ASP) - a declarative programming paradigm. We formulate the task of CCG parsing as a planning problem and use an ASP computational tool to compute solutions that correspond to valid parses. Compared to other approaches, there is no need to implement a specific parsing algorithm using such a declarative method. Our approach aims at producing all semantically distinct parse trees for a given sentence. From this goal, normalization and efficiency issues arise, and we deal with them by combining and extending existing strategies. We have implemented a CCG parsing tool kit - AspCcgTk - that uses ASP as its main computational means. The C&C supertagger can be used as a preprocessor within AspCcgTk, which allows us to achieve wide-coverage natural language parsing.