The Experts below are selected from a list of 434199 Experts worldwide ranked by ideXlab platform

Hinrich Schutze - One of the best experts on this subject based on the ideXlab platform.

  • quantifying the Contextualization of word representations with semantic class probing
    arXiv: Computation and Language, 2020
    Co-Authors: Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh, Hinrich Schutze
    Abstract:

    Pretrained language models have achieved a new state of the art on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the Contextualization of words in BERT. We quantify the amount of Contextualization, i.e., how well words are interpreted in context, by studying the extent to which semantic classes of a word can be inferred from its contextualized embeddings. Quantifying Contextualization helps in understanding and utilizing pretrained language models. We show that top layer representations achieve high accuracy inferring semantic classes; that the strongest Contextualization effects occur in the lower layers; that local context is mostly sufficient for semantic class inference; and that top layer representations are more task-specific after finetuning while lower layer representations are more transferable. Finetuning uncovers task related features, but pretrained knowledge is still largely preserved.

  • quantifying the Contextualization of word representations with semantic class probing
    Empirical Methods in Natural Language Processing, 2020
    Co-Authors: Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh, Hinrich Schutze
    Abstract:

    Pretrained language models achieve state-of-the-art results on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the Contextualization of words in BERT. We quantify the amount of Contextualization, i.e., how well words are interpreted in context, by studying the extent to which semantic classes of a word can be inferred from its contextualized embedding. Quantifying Contextualization helps in understanding and utilizing pretrained language models. We show that the top layer representations support highly accurate inference of semantic classes; that the strongest Contextualization effects occur in the lower layers; that local context is mostly sufficient for contextualizing words; and that top layer representations are more task-specific after finetuning while lower layer representations are more transferable. Finetuning uncovers task-related features, but pretrained knowledge about Contextualization is still well preserved.

Josiane Mothe - One of the best experts on this subject based on the ideXlab platform.

  • CLEF 2017 Microblog Cultural Contextualization Lab Overview
    Experimental IR Meets Multilinguality Multimodality and Interaction, 2017
    Co-Authors: Liana Ermakova, Philippe Mulhem, Lorraine Goeuriot, Josiane Mothe, Jian-yun Nie, Eric Sanjuan
    Abstract:

    MC2 CLEF 2017 lab deals with how cultural context of a microblog affects its social impact at large. This involves microblog search, classification, filtering, language recognition, localization, entity extraction, linking open data, and summarization. Regular Lab participants have access to the private massive multilingual microblog stream of The Festival Galleries project. Festivals have a large presence on social media. The resulting mircroblog stream and related URLs is appropriate to experiment advanced social media search and mining methods. A collection of 70,000,000 microblogs over 18 months dealing with cultural events in all languages has been released to test multilingual content analysis and microblog search. For content analysis topics were in any language and results were expected in four languages: English, Spanish, French, and Portuguese. For microblog search topics were in four languages: Arabic, English, French and Spanish, and results were expected in any language.

  • Tweet Data Mining: the Cultural Microblog Contextualization Data Set
    2016
    Co-Authors: Yassine Rkha Chaham, Clémentine Scohy, Sébastien Déjean, Josiane Mothe
    Abstract:

    This paper presents an overview of the data set that was used for the Cultural Microblog Contextualization Workshop at CLEF 2016 and more specifically for the task 1: tweet Contextualization. In this paper we first present a descriptive analysis of the data: we consider the variables or features associated with the tweets and analyse them. Then we also analyse the tweet textual content. The results of this work correspond to a first step toward data quality checking. It can also useful in order to understand better the data and its usefulness for some tasks or case studies.

  • Cultural micro-blog Contextualization 2016 Workshop Overview: data and pilot tasks
    2016
    Co-Authors: Liana Ermakova, Philippe Mulhem, Lorraine Goeuriot, Josiane Mothe, Jian-yun Nie, Eric Sanjuan
    Abstract:

    CLEF Cultural micro-blog Contextualization Workshop is aiming at providing the research community with data sets to gather, organize and deliver relevant social data related to events generating a large number of micro-blog posts and web documents. It is also devoted to discussing tasks to be run from this data set and that could serve applications.

  • Overview of the CLEF 2016 Cultural Micro-blog Contextualization Workshop
    2016
    Co-Authors: Lorraine Goeuriot, Philippe Mulhem, Josiane Mothe, Fionn Murtagh, Eric Sanjuan
    Abstract:

    CLEF Cultural micro-blog Contextualization Workshop is aiming at providing the research community with data sets to gather, organize and deliver relevant social data related to events generating a large number of micro-blog posts and web documents. It is also devoted to discussing tasks to be run from this data set and that could serve applications.

  • inex tweet Contextualization task
    Information Processing and Management, 2016
    Co-Authors: Patrice Bellot, Eric Sanjuan, Josiane Mothe, Véronique Moriceau, Xavier Tannier
    Abstract:

    A full summary report on the four-year long Tweet Contextualization task.A detail on evaluation metrics and framework we developed for tweet Contextualization evaluation.A deep analysis of what the participants suggested in their approaches by categorizing the various methods.A description of the data made available to the community. Microblogging platforms such as Twitter are increasingly used for on-line client and market analysis. This motivated the proposal of a new track at CLEF INEX lab of Tweet Contextualization. The objective of this task was to help a user to understand a tweet by providing him with a short explanatory summary (500 words). This summary should be built automatically using resources like Wikipedia and generated by extracting relevant passages and aggregating them into a coherent summary.Running for four years, results show that the best systems combine NLP techniques with more traditional methods. More precisely the best performing systems combine passage retrieval, sentence segmentation and scoring, named entity recognition, text part-of-speech (POS) analysis, anaphora detection, diversity content measure as well as sentence reordering.This paper provides a full summary report on the four-year long task. While yearly overviews focused on system results, in this paper we provide a detailed report on the approaches proposed by the participants and which can be considered as the state of the art for this task. As an important result from the 4 years competition, we also describe the open access resources that have been built and collected. The evaluation measures for automatic summarization designed in DUC or MUC were not appropriate to evaluate tweet Contextualization, we explain why and depict in detailed the LogSim measure used to evaluate informativeness of produced contexts or summaries. Finally, we also mention the lessons we learned and that it is worth considering when designing a task.

Mengjie Zhao - One of the best experts on this subject based on the ideXlab platform.

  • quantifying the Contextualization of word representations with semantic class probing
    arXiv: Computation and Language, 2020
    Co-Authors: Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh, Hinrich Schutze
    Abstract:

    Pretrained language models have achieved a new state of the art on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the Contextualization of words in BERT. We quantify the amount of Contextualization, i.e., how well words are interpreted in context, by studying the extent to which semantic classes of a word can be inferred from its contextualized embeddings. Quantifying Contextualization helps in understanding and utilizing pretrained language models. We show that top layer representations achieve high accuracy inferring semantic classes; that the strongest Contextualization effects occur in the lower layers; that local context is mostly sufficient for semantic class inference; and that top layer representations are more task-specific after finetuning while lower layer representations are more transferable. Finetuning uncovers task related features, but pretrained knowledge is still largely preserved.

  • quantifying the Contextualization of word representations with semantic class probing
    Empirical Methods in Natural Language Processing, 2020
    Co-Authors: Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh, Hinrich Schutze
    Abstract:

    Pretrained language models achieve state-of-the-art results on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the Contextualization of words in BERT. We quantify the amount of Contextualization, i.e., how well words are interpreted in context, by studying the extent to which semantic classes of a word can be inferred from its contextualized embedding. Quantifying Contextualization helps in understanding and utilizing pretrained language models. We show that the top layer representations support highly accurate inference of semantic classes; that the strongest Contextualization effects occur in the lower layers; that local context is mostly sufficient for contextualizing words; and that top layer representations are more task-specific after finetuning while lower layer representations are more transferable. Finetuning uncovers task-related features, but pretrained knowledge about Contextualization is still well preserved.

Maarten De Rijke - One of the best experts on this subject based on the ideXlab platform.

  • weakly supervised Contextualization of knowledge graph facts
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2018
    Co-Authors: Nikos Voskarides, Edgar Meij, Ridho Reinanda, Abhinav Khaitan, Miles Osborne, Giorgio Stefanoni, Prabhanjan Kambadur, Maarten De Rijke
    Abstract:

    Knowledge graphs (KGs) model facts about the world; they consist of nodes (entities such as companies and people) that are connected by edges (relations such as founderOf ). Facts encoded in KGs are frequently used by search applications to augment result pages. When presenting a KG fact to the user, providing other facts that are pertinent to that main fact can enrich the user experience and support exploratory information needs. \em KG fact Contextualization is the task of augmenting a given KG fact with additional and useful KG facts. The task is challenging because of the large size of KGs; discovering other relevant facts even in a small neighborhood of the given fact results in an enormous amount of candidates. We introduce a neural fact Contextualization method (\em NFCM ) to address the KG fact Contextualization task. NFCM first generates a set of candidate facts in the neighborhood of a given fact and then ranks the candidate facts using a supervised learning to rank model. The ranking model combines features that we automatically learn from data and that represent the query-candidate facts with a set of hand-crafted features we devised or adjusted for this task. In order to obtain the annotations required to train the learning to rank model at scale, we generate training data automatically using distant supervision on a large entity-tagged text corpus. We show that ranking functions learned on this data are effective at contextualizing KG facts. Evaluation using human assessors shows that it significantly outperforms several competitive baselines.

  • weakly supervised Contextualization of knowledge graph facts
    arXiv: Information Retrieval, 2018
    Co-Authors: Nikos Voskarides, Edgar Meij, Ridho Reinanda, Abhinav Khaitan, Miles Osborne, Giorgio Stefanoni, Prabhanjan Kambadur, Maarten De Rijke
    Abstract:

    Knowledge graphs (KGs) model facts about the world, they consist of nodes (entities such as companies and people) that are connected by edges (relations such as founderOf). Facts encoded in KGs are frequently used by search applications to augment result pages. When presenting a KG fact to the user, providing other facts that are pertinent to that main fact can enrich the user experience and support exploratory information needs. KG fact Contextualization is the task of augmenting a given KG fact with additional and useful KG facts. The task is challenging because of the large size of KGs, discovering other relevant facts even in a small neighborhood of the given fact results in an enormous amount of candidates. We introduce a neural fact Contextualization method (NFCM) to address the KG fact Contextualization task. NFCM first generates a set of candidate facts in the neighborhood of a given fact and then ranks the candidate facts using a supervised learning to rank model. The ranking model combines features that we automatically learn from data and that represent the query-candidate facts with a set of hand-crafted features we devised or adjusted for this task. In order to obtain the annotations required to train the learning to rank model at scale, we generate training data automatically using distant supervision on a large entity-tagged text corpus. We show that ranking functions learned on this data are effective at contextualizing KG facts. Evaluation using human assessors shows that it significantly outperforms several competitive baselines.

  • query dependent Contextualization of streaming data
    European Conference on Information Retrieval, 2014
    Co-Authors: Nikos Voskarides, Daan Odijk, Manos Tsagkias, Wouter Weerkamp, Maarten De Rijke
    Abstract:

    We propose a method for linking entities in a stream of short textual documents that takes into account context both inside a document and inside the history of documents seen so far. Our method uses a generic optimization framework for combining several entity ranking functions, and we introduce a global control function to control optimization. Our results demonstrate the effectiveness of combining entity ranking functions that take into account context, which is further boosted by 6% when we use an informed global control function.

Liana Ermakova - One of the best experts on this subject based on the ideXlab platform.

  • CLEF 2017 Microblog Cultural Contextualization Lab Overview
    Experimental IR Meets Multilinguality Multimodality and Interaction, 2017
    Co-Authors: Liana Ermakova, Philippe Mulhem, Lorraine Goeuriot, Josiane Mothe, Jian-yun Nie, Eric Sanjuan
    Abstract:

    MC2 CLEF 2017 lab deals with how cultural context of a microblog affects its social impact at large. This involves microblog search, classification, filtering, language recognition, localization, entity extraction, linking open data, and summarization. Regular Lab participants have access to the private massive multilingual microblog stream of The Festival Galleries project. Festivals have a large presence on social media. The resulting mircroblog stream and related URLs is appropriate to experiment advanced social media search and mining methods. A collection of 70,000,000 microblogs over 18 months dealing with cultural events in all languages has been released to test multilingual content analysis and microblog search. For content analysis topics were in any language and results were expected in four languages: English, Spanish, French, and Portuguese. For microblog search topics were in four languages: Arabic, English, French and Spanish, and results were expected in any language.

  • Cultural micro-blog Contextualization 2016 Workshop Overview: data and pilot tasks
    2016
    Co-Authors: Liana Ermakova, Philippe Mulhem, Lorraine Goeuriot, Josiane Mothe, Jian-yun Nie, Eric Sanjuan
    Abstract:

    CLEF Cultural micro-blog Contextualization Workshop is aiming at providing the research community with data sets to gather, organize and deliver relevant social data related to events generating a large number of micro-blog posts and web documents. It is also devoted to discussing tasks to be run from this data set and that could serve applications.

  • Short text Contextualization in information retrieval : application to tweet Contextualization and automatic query expansion
    2016
    Co-Authors: Liana Ermakova
    Abstract:

    The efficient communication tends to follow the principle of the least effort. According to this principle, using a given language interlocutors do not want to work any harder than necessary to reach understanding. This fact leads to the extreme compression of texts especially in electronic communication, e.g. microblogs, SMS, search queries. However, sometimes these texts are not self-contained and need to be explained since understanding them requires knowledge of terminology, named entities or related facts. The main goal of this research is to provide a context to a user or a system from a textual resource.The first aim of this work is to help a user to better understand a short message by extracting a context from an external source like a text collection, the Web or the Wikipedia by means of text summarization. To this end we developed an approach for automatic multi-document summarization and we applied it to short message Contextualization, in particular to tweet Contextualization. The proposed method is based on named entity recognition, part-of-speech weighting and sentence quality measuring. In contrast to previous research, we introduced an algorithm for smoothing from the local context. Our approach exploits topic-comment structure of a text. Moreover, we developed a graph-based algorithm for sentence reordering. The method has been evaluated at INEX/CLEF tweet Contextualization track. We provide the evaluation results over the 4 years of the track. The method was also adapted to snippet retrieval. The evaluation results indicate good performance of the approach.

  • A Method for Short Message Contextualization: Experiments at CLEF/INEX
    2015
    Co-Authors: Liana Ermakova
    Abstract:

    This paper presents the approach we developed for automatic multi-document summarization applied to short message Contextualization, in particular to tweet Contextualization. The proposed method is based on named entity recognition, part-of-speech weighting and sentence quality measuring. In contrast to previous research, we introduced an algorithm from smoothing from the local context. Our approach exploits topic-comment structure of a text. Moreover, we developed a graph-based algorithm for sentence reordering. The method has been evaluated at INEX/CLEF tweet Contextualization track. We provide the evaluation results over the 4 years of the track. The method was also adapted to snippet retrieval and query expansion. The evaluation results indicate good performance of the approach.

  • IRIT at INEX 2014 : Tweet Contextualization Track
    2014
    Co-Authors: Liana Ermakova, Josiane Mothe
    Abstract:

    The paper presents IRIT's approach used at INEX Tweet Contextualization Track 2014. Systems had to provide a context to a tweet from the perspective of the entity. This year we further modified our approach presented at INEX 2011, 2012 and 2013 underlain by the product of different measures based on smoothing from local context, named entity recognition, part-ofspeech weighting and sentence quality analysis. We introduced two ways to link an entity and a tweet, namely (1) concatenation of the entity and the tweet and (2) usage of the results obtained for the entity as a restriction to filter results retrieved for the tweet. Besides, we examined the influence of topic-comment relationship on Contextualization.