The Experts below are selected from a list of 13113 Experts worldwide ranked by ideXlab platform

Nigel Collier - One of the best experts on this subject based on the ideXlab platform.

  • bio medical Entity Extraction using support vector machines
    Artificial Intelligence in Medicine, 2005
    Co-Authors: Koichi Takeuchi, Nigel Collier
    Abstract:

    Objective:: Support vector machines (SVMs) have achieved state-of-the-art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named Entity task to special domains with large-scale terminologies such as those in medicine and related disciplines. Methods and materials:: The foundation for the model is a sample of text annotated by a domain expert according to an ontology of concepts, properties and relations. The model then learns to annotate unseen terms in new texts and contexts. The results can be used for a variety of intelligent language processing applications. We illustrate SVMs capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. Results:: Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance. Conclusion:: Our experiments indicate a relationship between feature window size and the amount of training data and that a combination of surface words, orthographic features and head noun features achieve the best performance among the feature sets tested.

  • bio medical Entity Extraction using support vector machines
    Meeting of the Association for Computational Linguistics, 2003
    Co-Authors: Koichi Takeuchi, Nigel Collier
    Abstract:

    Support Vector Machines have achieved state of the art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named Entity task to special domains with extensive terminologies such as those in medicine and related disciplines. We illustrate SVM's capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance.

Csaba Veres - One of the best experts on this subject based on the ideXlab platform.

  • named Entity Extraction for knowledge graphs a literature overview
    IEEE Access, 2020
    Co-Authors: Tareq Almoslmi, Marc Gallofre Ocana, Andreas L Opdahl, Csaba Veres
    Abstract:

    An enormous amount of digital information is expressed as natural-language (NL) text that is not easily processable by computers. Knowledge Graphs (KG) offer a widely used format for representing information in computer-processable form. Natural Language Processing (NLP) is therefore needed for mining (or lifting) knowledge graphs from NL texts. A central part of the problem is to extract the named entities in the text. The paper presents an overview of recent advances in this area, covering: Named Entity Recognition (NER), Named Entity Disambiguation (NED), and Named Entity Linking (NEL). We comment that many approaches to NED and NEL are based on older approaches to NER and need to leverage the outputs of state-of-the-art NER systems. There is also a need for standard methods to evaluate and compare named-Entity Extraction approaches. We observe that NEL has recently moved from being stepwise and isolated into an integrated process along two dimensions: the first is that previously sequential steps are now being integrated into end-to-end processes, and the second is that entities that were previously analysed in isolation are now being lifted in each other's context. The current culmination of these trends are the deep-learning approaches that have recently reported promising results.

Keun Young Kang - One of the best experts on this subject based on the ideXlab platform.

  • PKDE4J: Entity and relation Extraction for public knowledge discovery
    Journal of Biomedical Informatics, 2015
    Co-Authors: Min Song, Won Chul Kim, Dahee Lee, Go Eun Heo, Keun Young Kang
    Abstract:

    Due to an enormous number of scientific publications that cannot be handled manually, there is a rising interest in text-mining techniques for automated information Extraction, especially in the biomedical field. Such techniques provide effective means of information search, knowledge discovery, and hypothesis generation. Most previous studies have primarily focused on the design and performance improvement of either named Entity recognition or relation Extraction. In this paper, we present PKDE4J, a comprehensive text-mining system that integrates dictionary-based Entity Extraction and rule-based relation Extraction in a highly flexible and extensible framework. Starting with the Stanford CoreNLP, we developed the system to cope with multiple types of entities and relations. The system also has fairly good performance in terms of accuracy as well as the ability to configure text-processing components. We demonstrate its competitive performance by evaluating it on many corpora and found that it surpasses existing systems with average F-measures of 85% for Entity Extraction and 81% for relation Extraction.

William W Cohen - One of the best experts on this subject based on the ideXlab platform.

  • revisiting semi supervised learning with graph embeddings
    International Conference on Machine Learning, 2016
    Co-Authors: Zhilin Yang, William W Cohen, Ruslan Salakhutdinov
    Abstract:

    We present a semi-supervised learning framework based on graph embeddings. Given a graph between instances, we train an embedding for each instance to jointly predict the class label and the neighborhood context in the graph. We develop both transductive and inductive variants of our method. In the transductive variant of our method, the class labels are determined by both the learned embeddings and input feature vectors, while in the inductive variant, the embeddings are defined as a parametric function of the feature vectors, so predictions can be made on instances not seen during training. On a large and diverse set of benchmark tasks, including text classification, distantly supervised Entity Extraction, and Entity classification, we show improved performance over many of the existing models.

  • joint information Extraction and reasoning a scalable statistical relational learning approach
    International Joint Conference on Natural Language Processing, 2015
    Co-Authors: William Yang Wang, William W Cohen
    Abstract:

    A standard pipeline for statistical relational learning involves two steps: one first constructs the knowledge base (KB) from text, and then performs the learning and reasoning tasks using probabilistic first-order logics. However, a key issue is that information Extraction (IE) errors from text affect the quality of the KB, and propagate to the reasoning task. In this paper, we propose a statistical relational learning model for joint information Extraction and reasoning. More specifically, we incorporate context-based Entity Extraction with structure learning (SL) in a scalable probabilistic logic framework. We then propose a latent context invention (LCI) approach to improve the performance. In experiments, we show that our approach outperforms state-of-the-art baselines over three real-world Wikipedia datasets from multiple domains; that joint learning and inference for IE and SL significantly improve both tasks; that latent context invention further improves the results.

  • exploiting dictionaries in named Entity Extraction combining semi markov Extraction processes and data integration methods
    Knowledge Discovery and Data Mining, 2004
    Co-Authors: William W Cohen, Sunita Sarawagi
    Abstract:

    We consider the problem of improving named Entity recognition (NER) systems by using external dictionaries---more specifically, the problem of extending state-of-the-art NER systems by incorporating information about the similarity of extracted entities to entities in an external dictionary. This is difficult because most high-performance named Entity recognition systems operate by sequentially classifying words as to whether or not they participate in an Entity name; however, the most useful similarity measures score entire candidate names. To correct this mismatch we formalize a semi-Markov Extraction process, which is based on sequentially classifying segments of several adjacent words, rather than single words. In addition to allowing a natural way of coupling high-performance NER methods and high-performance similarity functions, this formalism also allows the direct use of other useful Entity-level features, and provides a more natural formulation of the NER problem than sequential word classification. Experiments in multiple domains show that the new model can substantially improve Extraction performance over previous methods for using external dictionaries in NER.

Marco Pennacchiotti - One of the best experts on this subject based on the ideXlab platform.

  • open Entity Extraction from web search query logs
    International Conference on Computational Linguistics, 2010
    Co-Authors: Alpa Jain, Marco Pennacchiotti
    Abstract:

    In this paper we propose a completely unsupervised method for open-domain Entity Extraction and clustering over query logs. The underlying hypothesis is that classes defined by mining search user activity may significantly differ from those typically considered over web documents, in that they better model the user space, i.e. users' perception and interests. We show that our method outperforms state of the art (semi-)supervised systems based either on web documents or on query logs (16% gain on the clustering task). We also report evidence that our method successfully supports a real world application, namely keyword generation for sponsored search.

  • Entity Extraction via ensemble semantics
    Empirical Methods in Natural Language Processing, 2009
    Co-Authors: Marco Pennacchiotti, Patrick Pantel
    Abstract:

    Combining information Extraction systems yields significantly higher quality resources than each system in isolation. In this paper, we generalize such a mixing of sources and features in a framework called Ensemble Semantics. We show very large gains in Entity Extraction by combining state-of-the-art distributional and pattern-based systems with a large set of features from a webcrawl, query logs, and Wikipedia. Experimental results on a web-scale Extraction of actors, athletes and musicians show significantly higher mean average precision scores (29% gain) compared with the current state of the art.