The Experts below are selected from a list of 269736 Experts worldwide ranked by ideXlab platform

Bruce W Croft - One of the best experts on this subject based on the ideXlab platform.

  • embedding based query language Models
    International Conference on the Theory of Information Retrieval, 2016
    Co-Authors: Hamed Zamani, Bruce W Croft
    Abstract:

    Word embeddings, which are low-dimensional vector representations of vocabulary terms that capture the semantic similarity between them, have recently been shown to achieve impressive performance in many natural language processing tasks. The use of word embeddings in information retrieval, however, has only begun to be studied. In this paper, we explore the use of word embeddings to enhance the accuracy of query language Models in the ad-hoc retrieval task. To this end, we propose to use word embeddings to incorporate and weight terms that do not occur in the query, but are semantically related to the query terms. We describe two embedding-based query expansion Models with different assumptions. Since pseudo-Relevance feedback methods that use the top retrieved documents to update the original query Model are well-known to be effective, we also develop an embedding-based Relevance Model, an extension of the effective and robust Relevance Model approach. In these Models, we transform the similarity values obtained by the widely-used cosine similarity with a sigmoid function to have more discriminative semantic similarity values. We evaluate our proposed methods using three TREC newswire and web collections. The experimental results demonstrate that the embedding-based methods significantly outperform competitive baselines in most cases. The embedding-based methods are also shown to be more robust than the baselines.

  • a deterministic resampling method using overlapping document clusters for pseudo Relevance feedback
    Information Processing and Management, 2013
    Co-Authors: Bruce W Croft
    Abstract:

    Typical pseudo-Relevance feedback methods assume the top-retrieved documents are relevant and use these pseudo-relevant documents to expand terms. The initial retrieval set can, however, contain a great deal of noise. In this paper, we present a cluster-based resampling method to select novel pseudo-relevant documents based on Lavrenko's Relevance Model approach. The main idea is to use overlapping clusters to find dominant documents for the initial retrieval set, and to repeatedly use these documents to emphasize the core topics of a query. The proposed resampling method can skip some documents in the initial high-ranked documents and deterministically construct overlapping clusters as sampling units. The hypothesis behind using overlapping clusters is that a good representative document for a query may have several nearest neighbors with high similarities, participating in several different clusters. Experimental results on large-scale web TREC collections show significant improvements over the baseline Relevance Model. To justify the proposed approach, we examine the Relevance density and redundancy ratio of feedback documents. A higher Relevance density will result in greater retrieval accuracy, ultimately approaching true Relevance feedback. The resampling approach shows higher Relevance density than the baseline Relevance Model on all collections, resulting in better retrieval accuracy in pseudo-Relevance feedback.

  • a field Relevance Model for structured document retrieval
    European Conference on Information Retrieval, 2012
    Co-Authors: Jinyoung Kim, Bruce W Croft
    Abstract:

    Many search applications involve documents with structure or fields. Since query terms often are related to specific structural components, mapping queries to fields and assigning weights to those fields is critical for retrieval effectiveness. Although several field-based retrieval Models have been developed, there has not been a formal justification of field weighting.In this work, we aim to improve the field weighting for structured document retrieval. We first introduce the notion of field Relevance as the generalization of field weights, and discuss how it can be estimated using relevant documents, which effectively implements Relevance feedback for field weighting. We then propose a framework for estimating field Relevance based on the combination of several sources. Evaluation on several structured document collections show that field weighting based on the suggested framework improves retrieval effectiveness significantly.

  • a unified Relevance Model for opinion retrieval
    Conference on Information and Knowledge Management, 2009
    Co-Authors: Xuanjing Huang, Bruce W Croft
    Abstract:

    Representing the information need is the greatest challenge for opinion retrieval. Typical queries for opinion retrieval are composed of either just content words, or content words with a small number of cue "opinion" words. Both are inadequate for retrieving opinionated documents. In this paper, we develop a general formal framework--the opinion Relevance Model--to represent an information need for opinion retrieval. We explore a series of methods to automatically identify the most appropriate opinion words for query expansion, including using query independent sentiment resources. We also propose a Relevance feedback-based approach to extract opinion words. Both query-independent and query-dependent methods can also be integrated into a more effective mixture Relevance Model. Finally, opinion retrieval experiments are presented for the Blog06 and COAE08 text collections. The results show that, significant improvements can always be obtained by this opinion Relevance Model whether sentiment resources are available or not.

  • a cluster based resampling method for pseudo Relevance feedback
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008
    Co-Authors: Bruce W Croft, James Allan
    Abstract:

    Typical pseudo-Relevance feedback methods assume the top-retrieved documents are relevant and use these pseudo-relevant documents to expand terms. The initial retrieval set can, however, contain a great deal of noise. In this paper, we present a cluster-based resampling method to select better pseudo-relevant documents based on the Relevance Model. The main idea is to use document clusters to find dominant documents for the initial retrieval set, and to repeatedly feed the documents to emphasize the core topics of a query. Experimental results on large-scale web TREC collections show significant improvements over the Relevance Model. For justification of the resampling approach, we examine Relevance density of feedback documents. A higher Relevance density will result in greater retrieval accuracy, ultimately approaching true Relevance feedback. The resampling approach shows higher Relevance density than the baseline Relevance Model on all collections, resulting in better retrieval accuracy in pseudo-Relevance feedback. This result indicates that the proposed method is effective for pseudo-Relevance feedback.

James Allan - One of the best experts on this subject based on the ideXlab platform.

  • A context-dependent Relevance Model
    Journal of the Association for Information Science and Technology, 2015
    Co-Authors: Edward Kai Fung Dang, Robert W. P. Luk, James Allan
    Abstract:

    Numerous past studies have demonstrated the effectiveness of the Relevance ModelRM for information retrieval IR. This approach enables Relevance or pseudo-Relevance feedback to be incorporated within the language Modeling framework of IR. In the traditional RM, the feedback information is used to improve the estimate of the query language Model. In this article, we introduce an extension of RM in the setting of Relevance feedback. Our method provides an additional way to incorporate feedback via the improvement of the document language Models. Specifically, we make use of the context information of known relevant and nonrelevant documents to obtain weighted counts of query terms for estimating the document language Models. The context information is based on the words unigrams or bigrams appearing within a text window centered on query terms. Experiments on several Text REtrieval Conference TREC collections show that our context-dependent Relevance Model can improve retrieval performance over the baseline RM. Together with previous studies within the BM25 framework, our current study demonstrates that the effectiveness of our method for using context information in IR is quite general and not limited to any specific retrieval Model.

  • a cluster based resampling method for pseudo Relevance feedback
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008
    Co-Authors: Bruce W Croft, James Allan
    Abstract:

    Typical pseudo-Relevance feedback methods assume the top-retrieved documents are relevant and use these pseudo-relevant documents to expand terms. The initial retrieval set can, however, contain a great deal of noise. In this paper, we present a cluster-based resampling method to select better pseudo-relevant documents based on the Relevance Model. The main idea is to use document clusters to find dominant documents for the initial retrieval set, and to repeatedly feed the documents to emphasize the core topics of a query. Experimental results on large-scale web TREC collections show significant improvements over the Relevance Model. For justification of the resampling approach, we examine Relevance density of feedback documents. A higher Relevance density will result in greater retrieval accuracy, ultimately approaching true Relevance feedback. The resampling approach shows higher Relevance density than the baseline Relevance Model on all collections, resulting in better retrieval accuracy in pseudo-Relevance feedback. This result indicates that the proposed method is effective for pseudo-Relevance feedback.

  • SIGIR - A cluster-based resampling method for pseudo-Relevance feedback
    Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '08, 2008
    Co-Authors: W. Bruce Croft, James Allan
    Abstract:

    Typical pseudo-Relevance feedback methods assume the top-retrieved documents are relevant and use these pseudo-relevant documents to expand terms. The initial retrieval set can, however, contain a great deal of noise. In this paper, we present a cluster-based resampling method to select better pseudo-relevant documents based on the Relevance Model. The main idea is to use document clusters to find dominant documents for the initial retrieval set, and to repeatedly feed the documents to emphasize the core topics of a query. Experimental results on large-scale web TREC collections show significant improvements over the Relevance Model. For justification of the resampling approach, we examine Relevance density of feedback documents. A higher Relevance density will result in greater retrieval accuracy, ultimately approaching true Relevance feedback. The resampling approach shows higher Relevance density than the baseline Relevance Model on all collections, resulting in better retrieval accuracy in pseudo-Relevance feedback. This result indicates that the proposed method is effective for pseudo-Relevance feedback.

Chengxiang Zhai - One of the best experts on this subject based on the ideXlab platform.

  • positional Relevance Model for pseudo Relevance feedback
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2010
    Co-Authors: Chengxiang Zhai
    Abstract:

    Pseudo-Relevance feedback is an effective technique for improving retrieval results. Traditional feedback algorithms use a whole feedback document as a unit to extract words for query expansion, which is not optimal as a document may cover several different topics and thus contain much irrelevant information. In this paper, we study how to effectively select from feedback documents those words that are focused on the query topic based on positions of terms in feedback documents. We propose a positional Relevance Model (PRM) to address this problem in a unified probabilistic way. The proposed PRM is an extension of the Relevance Model to exploit term positions and proximity so as to assign more weights to words closer to query words based on the intuition that words closer to query words are more likely to be related to the query topic. We develop two methods to estimate PRM based on different sampling processes. Experiment results on two large retrieval datasets show that the proposed PRM is effective and robust for pseudo-Relevance feedback, significantly outperforming the Relevance Model in both document-based feedback and passage-based feedback.

  • SIGIR - Positional Relevance Model for pseudo-Relevance feedback
    Proceeding of the 33rd international ACM SIGIR conference on Research and development in information retrieval - SIGIR '10, 2010
    Co-Authors: Chengxiang Zhai
    Abstract:

    Pseudo-Relevance feedback is an effective technique for improving retrieval results. Traditional feedback algorithms use a whole feedback document as a unit to extract words for query expansion, which is not optimal as a document may cover several different topics and thus contain much irrelevant information. In this paper, we study how to effectively select from feedback documents those words that are focused on the query topic based on positions of terms in feedback documents. We propose a positional Relevance Model (PRM) to address this problem in a unified probabilistic way. The proposed PRM is an extension of the Relevance Model to exploit term positions and proximity so as to assign more weights to words closer to query words based on the intuition that words closer to query words are more likely to be related to the query topic. We develop two methods to estimate PRM based on different sampling processes. Experiment results on two large retrieval datasets show that the proposed PRM is effective and robust for pseudo-Relevance feedback, significantly outperforming the Relevance Model in both document-based feedback and passage-based feedback.

  • a study of term proximity and document weighting normalization in pseudo Relevance feedback uiuc at trec 2009 million query track
    Text REtrieval Conference, 2009
    Co-Authors: V Vinod G Vydiswaran, Kavita Ganesan, Chengxiang Zhai
    Abstract:

    In this paper, we report our experiments in the TREC 2009 Million Query Track. Our flrst line of study is on proximitybased feedback, in which we propose a positional Relevance Model (PRM) to exploit term proximity evidence so as to assign more weights to expansion words that are closer to query words in feedback documents. The second line of study is to improve the weighting of feedback documents in the Relevance Model by using a regression-based method to approximate the probability of Relevance (and thus the name RegRM). In the third line of study, we test a supervised approach for query classiflcation. Besides, we also evaluate a selective pseudo feedback strategy which stops pseudo feedback for precision-oriented queries and only uses it for recall-oriented ones. The proposed PRM has shown clear improvements over the Relevance Model for pseudo feedback, suggesting that capturing the term proximity heuristic appropriately could lead to a better feedback Model. RegRM performs as well as Relevance Model, but no noticeable improvement is observed. Unfortunately, the proposed query classiflcation methods appear to not work well. The results also show that the proposed selective pseudo feedback may not work well, since precision-oriented queries can also beneflt from pseudo feedback, though not as much as recall-oriented queries.

  • a comparative study of methods for estimating query language Models with pseudo feedback
    Conference on Information and Knowledge Management, 2009
    Co-Authors: Yuanhua Lv, Chengxiang Zhai
    Abstract:

    We systematically compare five representative state-of-the-art methods for estimating query language Models with pseudo feedback in ad hoc information retrieval, including two variants of the Relevance language Model, two variants of the mixture feedback Model, and the divergence minimization estimation method. Our experiment results show that a variant of Relevance Model and a variant of the mixture Model tend to outperform other methods. We further propose several heuristics that are intuitively related to the good retrieval performance of an estimation method, and show that the variations in how these heuristics are implemented in different methods provide a good explanation of many empirical observations.

  • probabilistic Relevance Models based on document and query generation
    2003
    Co-Authors: John Lafferty, Chengxiang Zhai
    Abstract:

    We give a unified account of the probabilistic semantics underlying the language Modeling approach and the traditional probabilistic Model for information retrieval, showing that the two approaches can be viewed as being equivalent probabilistically, since they are based on different factorizations of the same generative Relevance Model. We also discuss how the two approaches lead to different retrieval frameworks in practice, since they involve component Models that are estimated quite differently.

W. Bruce Croft - One of the best experts on this subject based on the ideXlab platform.

  • Query structuring and expansion with two-stage term dependence for Japanese web retrieval
    Information Retrieval, 2009
    Co-Authors: Koji Eguchi, W. Bruce Croft
    Abstract:

    In this paper, we propose a new term dependence Model for information retrieval, which is based on a theoretical framework using Markov random fields. We assume two types of dependencies of terms given in a query: (i) long-range dependencies that may appear for instance within a passage or a sentence in a target document, and (ii) short-range dependencies that may appear for instance within a compound word in a target document. Based on this assumption, our two-stage term dependence Model captures both long-range and short-range term dependencies differently, when more than one compound word appear in a query. We also investigate how query structuring with term dependence can improve the performance of query expansion using a Relevance Model. The Relevance Model is constructed using the retrieval results of the structured query with term dependence to expand the query. We show that our term dependence Model works well, particularly when using query structuring with compound words, through experiments using a 100-gigabyte test collection of web documents mostly written in Japanese. We also show that the performance of the Relevance Model can be significantly improved by using the structured query with our term dependence Model.

  • CIKM - A unified Relevance Model for opinion retrieval
    Proceeding of the 18th ACM conference on Information and knowledge management - CIKM '09, 2009
    Co-Authors: Xuanjing Huang, W. Bruce Croft
    Abstract:

    Representing the information need is the greatest challenge for opinion retrieval. Typical queries for opinion retrieval are composed of either just content words, or content words with a small number of cue "opinion" words. Both are inadequate for retrieving opinionated documents. In this paper, we develop a general formal framework--the opinion Relevance Model--to represent an information need for opinion retrieval. We explore a series of methods to automatically identify the most appropriate opinion words for query expansion, including using query independent sentiment resources. We also propose a Relevance feedback-based approach to extract opinion words. Both query-independent and query-dependent methods can also be integrated into a more effective mixture Relevance Model. Finally, opinion retrieval experiments are presented for the Blog06 and COAE08 text collections. The results show that, significant improvements can always be obtained by this opinion Relevance Model whether sentiment resources are available or not.

  • SIGIR - A cluster-based resampling method for pseudo-Relevance feedback
    Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '08, 2008
    Co-Authors: W. Bruce Croft, James Allan
    Abstract:

    Typical pseudo-Relevance feedback methods assume the top-retrieved documents are relevant and use these pseudo-relevant documents to expand terms. The initial retrieval set can, however, contain a great deal of noise. In this paper, we present a cluster-based resampling method to select better pseudo-relevant documents based on the Relevance Model. The main idea is to use document clusters to find dominant documents for the initial retrieval set, and to repeatedly feed the documents to emphasize the core topics of a query. Experimental results on large-scale web TREC collections show significant improvements over the Relevance Model. For justification of the resampling approach, we examine Relevance density of feedback documents. A higher Relevance density will result in greater retrieval accuracy, ultimately approaching true Relevance feedback. The resampling approach shows higher Relevance density than the baseline Relevance Model on all collections, resulting in better retrieval accuracy in pseudo-Relevance feedback. This result indicates that the proposed method is effective for pseudo-Relevance feedback.

  • Relevance Models in Information Retrieval
    Language Modeling for Information Retrieval, 2003
    Co-Authors: Victor Lavrenko, W. Bruce Croft
    Abstract:

    We develop a simple statistical Model, called a Relevance Model, for capturing the notion of topical Relevance in information retrieval. Estimating probabilities of Relevance has been an important part of many previous retrieval Models, but we show how this estimation can be done in a more principled way based on a generative or language Model approach. In particular, we focus on estimating Relevance Models when training examples (examples of relevant documents) are not available. We describe extensive evaluations of the Relevance Model approach on the TREC ad-hoc retrieval and cross-language tasks. In both cases, rankings based on Relevance Models significantly outperform strong baseline approaches.

Doron Cohen - One of the best experts on this subject based on the ideXlab platform.

  • SIGIR - An Extended Relevance Model for Session Search
    Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2017
    Co-Authors: Nir Levine, Haggai Roitman, Doron Cohen
    Abstract:

    The session search task aims at best serving the user's information need given her previous search behavior during the session. We propose an extended Relevance Model that captures the user's dynamic information need in the session. Our Relevance Modelling approach is directly driven by the user's query reformulation (change) decisions and the estimate of how much the user's search behavior affects such decisions. Overall, we demonstrate that, the proposed approach significantly boosts session search performance.

  • An Extended Relevance Model for Session Search
    arXiv: Information Retrieval, 2017
    Co-Authors: Nir Levine, Haggai Roitman, Doron Cohen
    Abstract:

    The session search task aims at best serving the user's information need given her previous search behavior during the session. We propose an extended Relevance Model that captures the user's dynamic information need in the session. Our Relevance Modelling approach is directly driven by the user's query reformulation (change) decisions and the estimate of how much the user's search behavior affects such decisions. Overall, we demonstrate that, the proposed approach significantly boosts session search performance.