The Experts below are selected from a list of 84 Experts worldwide ranked by ideXlab platform
Rosa Stern - One of the best experts on this subject based on the ideXlab platform.
-
Population of a Knowledge Base for News Metadata from Unstructured Text and Web Data
2012Co-Authors: Rosa Stern, Benoît SagotAbstract:We present a practical use case of knowl- edge base (KB) population at the French news agency AFP. The target KB instances are en- tities relevant for news production and con- tent enrichment. In order to acquire uniquely identified entities over news wires, i.e. tex- tual data, and integrate the resulting KB in the Linked Data framework, a series of data mod- els need to be aligned: Web data resources are harvested for creating a wide coverage Entity Database, which is in turn used to link entities to their mentions in French news wires. Fi- nally, the extracted entities are selected for in- stantiation in the target KB. We describe our methodology along with the resources created and used for the target KB population.
-
Aleda, a free large-scale Entity Database for French
2012Co-Authors: Benoît Sagot, Rosa SternAbstract:Named Entity recognition, which focuses on the identification of the span and type of named Entity mentions in texts, has drawn the attention of the NLP community for a long time. However, many real-life applications need to know which real Entity each mention refers to. For such a purpose, often refered to as Entity resolution and linking, an inventory of entities is required in order to constitute a reference. In this paper, we describe how we extracted such a resource for French from freely available resources (the French Wikipedia and the GeoNames Database). We describe the results of an instrinsic evaluation of the resulting Entity Database, named Aleda, as well as those of a task-based evaluation in the context of a named Entity detection system. We also compare it with the NLGbAse Database (Charton and Torres-Moreno, 2010), a resource with similar objectives.
-
LREC - Aleda, a free large-scale Entity Database for French
2012Co-Authors: Benoît Sagot, Rosa SternAbstract:Named Entity recognition, which focuses on the identification of the span and type of named Entity mentions in texts, has drawn the attention of the NLP community for a long time. However, many real-life applications need to know which real Entity each mention refers to. For such a purpose, often refered to as Entity resolution and linking, an inventory of entities is required in order to constitute a reference. In this paper, we describe how we extracted such a resource for French from freely available resources (the French Wikipedia and the GeoNames Database). We describe the results of an instrinsic evaluation of the resulting Entity Database, named Aleda, as well as those of a task-based evaluation in the context of a named Entity detection system. We also compare it with the NLGbAse Database (Charton and Torres-Moreno, 2010), a resource with similar objectives.
-
AKBC-WEKEX@NAACL-HLT - Population of a Knowledge Base for News Metadata from Unstructured Text and Web Data
2012Co-Authors: Rosa Stern, Benoît SagotAbstract:We present a practical use case of knowledge base (KB) population at the French news agency AFP. The target KB instances are entities relevant for news production and content enrichment. In order to acquire uniquely identified entities over news wires, i.e. textual data, and integrate the resulting KB in the Linked Data framework, a series of data models need to be aligned: Web data resources are harvested for creating a wide coverage Entity Database, which is in turn used to link entities to their mentions in French news wires. Finally, the extracted entities are selected for instantiation in the target KB. We describe our methodology along with the resources created and used for the target KB population.
Benoît Sagot - One of the best experts on this subject based on the ideXlab platform.
-
Population of a Knowledge Base for News Metadata from Unstructured Text and Web Data
2012Co-Authors: Rosa Stern, Benoît SagotAbstract:We present a practical use case of knowl- edge base (KB) population at the French news agency AFP. The target KB instances are en- tities relevant for news production and con- tent enrichment. In order to acquire uniquely identified entities over news wires, i.e. tex- tual data, and integrate the resulting KB in the Linked Data framework, a series of data mod- els need to be aligned: Web data resources are harvested for creating a wide coverage Entity Database, which is in turn used to link entities to their mentions in French news wires. Fi- nally, the extracted entities are selected for in- stantiation in the target KB. We describe our methodology along with the resources created and used for the target KB population.
-
Aleda, a free large-scale Entity Database for French
2012Co-Authors: Benoît Sagot, Rosa SternAbstract:Named Entity recognition, which focuses on the identification of the span and type of named Entity mentions in texts, has drawn the attention of the NLP community for a long time. However, many real-life applications need to know which real Entity each mention refers to. For such a purpose, often refered to as Entity resolution and linking, an inventory of entities is required in order to constitute a reference. In this paper, we describe how we extracted such a resource for French from freely available resources (the French Wikipedia and the GeoNames Database). We describe the results of an instrinsic evaluation of the resulting Entity Database, named Aleda, as well as those of a task-based evaluation in the context of a named Entity detection system. We also compare it with the NLGbAse Database (Charton and Torres-Moreno, 2010), a resource with similar objectives.
-
LREC - Aleda, a free large-scale Entity Database for French
2012Co-Authors: Benoît Sagot, Rosa SternAbstract:Named Entity recognition, which focuses on the identification of the span and type of named Entity mentions in texts, has drawn the attention of the NLP community for a long time. However, many real-life applications need to know which real Entity each mention refers to. For such a purpose, often refered to as Entity resolution and linking, an inventory of entities is required in order to constitute a reference. In this paper, we describe how we extracted such a resource for French from freely available resources (the French Wikipedia and the GeoNames Database). We describe the results of an instrinsic evaluation of the resulting Entity Database, named Aleda, as well as those of a task-based evaluation in the context of a named Entity detection system. We also compare it with the NLGbAse Database (Charton and Torres-Moreno, 2010), a resource with similar objectives.
-
AKBC-WEKEX@NAACL-HLT - Population of a Knowledge Base for News Metadata from Unstructured Text and Web Data
2012Co-Authors: Rosa Stern, Benoît SagotAbstract:We present a practical use case of knowledge base (KB) population at the French news agency AFP. The target KB instances are entities relevant for news production and content enrichment. In order to acquire uniquely identified entities over news wires, i.e. textual data, and integrate the resulting KB in the Linked Data framework, a series of data models need to be aligned: Web data resources are harvested for creating a wide coverage Entity Database, which is in turn used to link entities to their mentions in French news wires. Finally, the extracted entities are selected for instantiation in the target KB. We describe our methodology along with the resources created and used for the target KB population.
Eduardo Torres-schumann - One of the best experts on this subject based on the ideXlab platform.
-
KES (3) - Integrated document browsing and data acquisition for building large ontologies
Lecture Notes in Computer Science, 2006Co-Authors: Felix Weigel, Klaus U. Schulz, Levin Brunner, Eduardo Torres-schumannAbstract:Named entities (e.g., “Kofi Annan”, “Coca-Cola”, “Second World War”) are ubiquitous in web pages and other types of document and often provide a simplified picture of the document's content. We present an ontology currently containing 31,000 named entities in different languages from various domains such as history, geography, politics, sports, arts, etc., which is being developed at the University of Munich (LMU). The underlying graph data model is simple and yet extremely versatile in different application scenarios. We demonstrate a prototype of a graphical interface to both the ontology and to documents on the web or in a local document repository, with a tight interaction in both directions. Occurrences of concepts from the ontology are highlighted and hyperlinked in the documents. Unrecognized entities could be added to the Database and related to other concepts in a semiautomatic process. The Entity Database can also be used for extending full-text queries on the web or the repository to semantically close documents, and for indexing different kinds of named entities in the document repository. Similar to a programming IDE, the system illustrates how integrated browsing, search and update functionality contributes to the construction of high-quality ontologies, fundamental to the vision of a truly “semantic” web.
Felix Weigel - One of the best experts on this subject based on the ideXlab platform.
-
KES (3) - Integrated document browsing and data acquisition for building large ontologies
Lecture Notes in Computer Science, 2006Co-Authors: Felix Weigel, Klaus U. Schulz, Levin Brunner, Eduardo Torres-schumannAbstract:Named entities (e.g., “Kofi Annan”, “Coca-Cola”, “Second World War”) are ubiquitous in web pages and other types of document and often provide a simplified picture of the document's content. We present an ontology currently containing 31,000 named entities in different languages from various domains such as history, geography, politics, sports, arts, etc., which is being developed at the University of Munich (LMU). The underlying graph data model is simple and yet extremely versatile in different application scenarios. We demonstrate a prototype of a graphical interface to both the ontology and to documents on the web or in a local document repository, with a tight interaction in both directions. Occurrences of concepts from the ontology are highlighted and hyperlinked in the documents. Unrecognized entities could be added to the Database and related to other concepts in a semiautomatic process. The Entity Database can also be used for extending full-text queries on the web or the repository to semantically close documents, and for indexing different kinds of named entities in the document repository. Similar to a programming IDE, the system illustrates how integrated browsing, search and update functionality contributes to the construction of high-quality ontologies, fundamental to the vision of a truly “semantic” web.
Ariel Fuxman - One of the best experts on this subject based on the ideXlab platform.
-
ACL - Jigs and Lures: Associating Web Queries with Structured Entities
2011Co-Authors: Patrick Pantel, Ariel FuxmanAbstract:We propose methods for estimating the probability that an Entity from an Entity Database is associated with a web search query. Association is modeled using a query Entity click graph, blending general query click logs with vertical query click logs. Smoothing techniques are proposed to address the inherent data sparsity in such graphs, including interpolation using a query synonymy model. A large-scale empirical analysis of the smoothing techniques, over a 2-year click graph collected from a commercial search engine, shows significant reductions in modeling error. The association models are then applied to the task of recommending products to web queries, by annotating queries with products from a large catalog and then mining query-product associations through web search session analysis. Experimental analysis shows that our smoothing techniques improve coverage while keeping precision stable, and overall, that our top-performing model affects 9% of general web queries with 94% precision.
-
Proceedings of Association for Computational Linguistics - Human Language Technology (ACL-HLT-11) - Jigs and Lures: Associating Web Queries with Strongly-Typed Entities
2011Co-Authors: Patrick Pantel, Ariel FuxmanAbstract:We propose methods for estimating the probability that an Entity from an Entity Database is associated with a web search query. Association is modeled using a query Entity click graph, blending general query click logs with vertical query click logs. Smoothing techniques are proposed to address the inherent data sparsity in such graphs, including interpolation using a query synonymy model. A large-scale empirical analysis of the smoothing techniques, over a 2-year click graph collected from a commercial search engine, shows significant reductions in modeling error. The association models are then applied to the task of recommending products to web queries, by annotating queries with products from a large catalog and then mining queryproduct associations through web search session analysis. Experimental analysis shows that our smoothing techniques improve coverage while keeping precision stable, and overall, that our top-performing model affects 9% of general web queries with 94% precision.