The Experts below are selected from a list of 291 Experts worldwide ranked by ideXlab platform

Gerhard Weikum - One of the best experts on this subject based on the ideXlab platform.

  • Approximate Information filtering in peer to peer networks
    Web Information Systems Engineering, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Klaus Berberich, Manolis Koubarakis, Gerhard Weikum
    Abstract:

    Most approaches to Information filtering taken so far have the underlying hypothesis of potentially delivering notifications from every Information producer to subscribers. This exact publish/subscribe model creates an efficiency and scalability bottleneck, and might not even be desirable in certain applications. The work presented here puts forward MAPS, a novel approach to support Approximate Information filtering in a peer-to-peer environment. In MAPS a user subscribes to and monitors only carefully selected data sources, and receives notifications about interesting events from these sources only. This way scalability is enhanced by trading recall for lower message traffic. We define the protocols of a peer-to-peer architecture especially designed for Approximate Information filtering, and introduce new node selection strategies based on time series analysis techniques to improve data source selection. Our experimental evaluation shows that MAPS is scalable; it achieves high recall by monitoring only few data sources.

  • exploiting correlated keywords to improve Approximate Information filtering
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    Information filtering, also referred to as publish/subscribe, complements one-time searching since users are able to subscribe to Information sources and be notified whenever new documents of interest are published. In Approximate Information filtering only selected Information sources, that are likely to publish documents relevant to the user interests in the future, are monitored. To achieve this functionality, a subscriber exploits statistical metadata to identify promising publishers and index its continuous query only in those publishers. The statistics are maintained in a directory, usually on a per-keyword basis, thus disregarding possible correlations among keywords. Using this coarse Information, poor publisher selection may lead to poor filtering performance and thus loss of interesting documents.1 Based on the above observation, this work extends query routing techniques from the domain of distributed Information retrieval in peer-to-peer (P2P) networks, and provides new algorithms for exploiting the correlation among keywords in a filtering setting. We develop and evaluate two algorithms based on single-key and multi-key statistics and utilize two different synopses (Hash Sketches and KMV synopses) to compactly represent publishers. Our experimental evaluation using two real-life corpora with web and blog data demonstrates the filtering effectiveness of both approaches and highlights the different tradeoffs.

  • SIGIR - Exploiting correlated keywords to improve Approximate Information filtering
    Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '08, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    Information filtering, also referred to as publish/subscribe, complements one-time searching since users are able to subscribe to Information sources and be notified whenever new documents of interest are published. In Approximate Information filtering only selected Information sources, that are likely to publish documents relevant to the user interests in the future, are monitored. To achieve this functionality, a subscriber exploits statistical metadata to identify promising publishers and index its continuous query only in those publishers. The statistics are maintained in a directory, usually on a per-keyword basis, thus disregarding possible correlations among keywords. Using this coarse Information, poor publisher selection may lead to poor filtering performance and thus loss of interesting documents.1 Based on the above observation, this work extends query routing techniques from the domain of distributed Information retrieval in peer-to-peer (P2P) networks, and provides new algorithms for exploiting the correlation among keywords in a filtering setting. We develop and evaluate two algorithms based on single-key and multi-key statistics and utilize two different synopses (Hash Sketches and KMV synopses) to compactly represent publishers. Our experimental evaluation using two real-life corpora with web and blog data demonstrates the filtering effectiveness of both approaches and highlights the different tradeoffs.

  • efficient search and Approximate Information filtering in a distributed peer to peer environment of digital libraries
    ACM international conference on Digital libraries, 2007
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    We present a new architecture for efficient search and Approximate Information filtering in a distributed Peer-to-Peer (P2P) environment of Digital Libraries. The MinervaLight search system uses P2P techniques over a structured overlay network to distribute and maintain a directory of peer statistics. Based on the same directory, the MAPS Information filtering system provides an Approximate publish/subscribe functionality by monitoring the most promising digital libraries for publishing appropriate documents regarding a continuous query. In this paper, we discuss our system architecture that combines searching and Information filtering abilities. We show the system components of MinervaLight and explain the different facets of an Approximate pub/sub system for subscriptions that is high scalable, efficient, and notifies the subscribers about the most interesting publications in the P2P network of digital libraries. We also compare both approaches in terms of common properties and differences to show an overview of search and pub/sub using the same infrastructure.

  • ECDL - MinervaDL: an architecture for Information retrieval and filtering in distributed digital libraries
    Research and Advanced Technology for Digital Libraries, 2007
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    We present MinervaDL, a digital library architecture that supports Approximate Information retrieval and filtering functionality under a single unifying framework. The architecture of MinervaDL is based on the peer-to-peer search engine Minerva, and is able to handle huge amounts of data provided by digital libraries in a distributed and self-organizing way. The two-tier architecture and the use of the distributed hash table as the routing substrate provides an infrastructure for creating large networks of digital libraries with minimal administration costs. We discuss the main components of this architecture, present the protocols that regulate node interactions, and experimentally evaluate our approach.

Christian Zimmer - One of the best experts on this subject based on the ideXlab platform.

  • Approximate Information filtering in peer to peer networks
    Web Information Systems Engineering, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Klaus Berberich, Manolis Koubarakis, Gerhard Weikum
    Abstract:

    Most approaches to Information filtering taken so far have the underlying hypothesis of potentially delivering notifications from every Information producer to subscribers. This exact publish/subscribe model creates an efficiency and scalability bottleneck, and might not even be desirable in certain applications. The work presented here puts forward MAPS, a novel approach to support Approximate Information filtering in a peer-to-peer environment. In MAPS a user subscribes to and monitors only carefully selected data sources, and receives notifications about interesting events from these sources only. This way scalability is enhanced by trading recall for lower message traffic. We define the protocols of a peer-to-peer architecture especially designed for Approximate Information filtering, and introduce new node selection strategies based on time series analysis techniques to improve data source selection. Our experimental evaluation shows that MAPS is scalable; it achieves high recall by monitoring only few data sources.

  • exploiting correlated keywords to improve Approximate Information filtering
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    Information filtering, also referred to as publish/subscribe, complements one-time searching since users are able to subscribe to Information sources and be notified whenever new documents of interest are published. In Approximate Information filtering only selected Information sources, that are likely to publish documents relevant to the user interests in the future, are monitored. To achieve this functionality, a subscriber exploits statistical metadata to identify promising publishers and index its continuous query only in those publishers. The statistics are maintained in a directory, usually on a per-keyword basis, thus disregarding possible correlations among keywords. Using this coarse Information, poor publisher selection may lead to poor filtering performance and thus loss of interesting documents.1 Based on the above observation, this work extends query routing techniques from the domain of distributed Information retrieval in peer-to-peer (P2P) networks, and provides new algorithms for exploiting the correlation among keywords in a filtering setting. We develop and evaluate two algorithms based on single-key and multi-key statistics and utilize two different synopses (Hash Sketches and KMV synopses) to compactly represent publishers. Our experimental evaluation using two real-life corpora with web and blog data demonstrates the filtering effectiveness of both approaches and highlights the different tradeoffs.

  • Approximate Information Filtering in Structured Peer-to-Peer Networks
    2008
    Co-Authors: Christian Zimmer
    Abstract:

    Today';s content providers are naturally distributed and produce large amounts of Information every day, making peer-to-peer data management a promising approach offering scalability, adaptivity to dynamics, and failure resilience. In such systems, subscribing with a continuous query is of equal importance as one-time querying since it allows the user to cope with the high rate of Information production and avoid the cognitive overload of repeated searches. In the Information filtering setting users specify continuous queries, thus subscribing to newly appearing documents satisfying the query conditions. Contrary to existing approaches providing exact Information filtering functionality, this doctoral thesis introduces the concept of Approximate Information filtering, where users subscribe to only a few selected sources most likely to satisfy their Information demand. This way, efficiency and scalability are enhanced by trading a small reduction in recall for lower message traffic. This thesis contains the following contributions: (i) the first architecture to support Approximate Information filtering in structured peer-to-peer networks, (ii) novel strategies to select the most appropriate publishers by taking into account correlations among keywords, (iii) a prototype implementation for Approximate Information retrieval and filtering, and (iv) a digital library use case to demonstrate the integration of retrieval and filtering in a unified system. Heutige Content-Anbieter sind verteilt und produzieren riesige Mengen an Daten jeden Tag. Daher wird die Datenhaltung in Peer-to-Peer Netzen zu einem vielversprechenden Ansatz, der Skalierbarkeit, Anpassbarkeit an Dynamik und Ausfallsicherheit bietet. Fur solche Systeme besitzt das Abonnieren mit Daueranfragen die gleiche Wichtigkeit wie einmalige Anfragen, da dies dem Nutzer erlaubt, mit der hohen Datenrate umzugehen und gleichzeitig die Uberlastung durch erneutes Suchen verhindert. Im Information Filtering Szenario legen Nutzer Daueranfragen fest und abonnieren dadurch neue Dokumente, die die Anfrage erfullen. Im Gegensatz zu vorhandenen Ansatzen fur exaktes Information Filtering fuhrt diese Doktorarbeit das Konzept von approximativem Information Filtering ein. Ein Nutzer abonniert nur wenige ausgewahlte Quellen, die am ehesten die Anfrage erfullen werden. Effizienz und Skalierbarkeit werden verbessert, indem Recall gegen einen geringeren Nachrichtenverkehr eingetauscht wird. Diese Arbeit beinhaltet folgende Beitrage: (i) die erste Architektur fur approximatives Information Filtering in strukturierten Peer-to-Peer Netzen, (ii) Strategien zur Wahl der besten Anbieter unter Berucksichtigung von Schlusselworter-Korrelationen, (iii) ein Prototyp, der approximatives Information Retrieval und Filtering realisiert und (iv) ein Anwendungsfall fur Digitale Bibliotheken, der beide Funktionalitaten in einem vereinten System aufzeigt.

  • SIGIR - Exploiting correlated keywords to improve Approximate Information filtering
    Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '08, 2008
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    Information filtering, also referred to as publish/subscribe, complements one-time searching since users are able to subscribe to Information sources and be notified whenever new documents of interest are published. In Approximate Information filtering only selected Information sources, that are likely to publish documents relevant to the user interests in the future, are monitored. To achieve this functionality, a subscriber exploits statistical metadata to identify promising publishers and index its continuous query only in those publishers. The statistics are maintained in a directory, usually on a per-keyword basis, thus disregarding possible correlations among keywords. Using this coarse Information, poor publisher selection may lead to poor filtering performance and thus loss of interesting documents.1 Based on the above observation, this work extends query routing techniques from the domain of distributed Information retrieval in peer-to-peer (P2P) networks, and provides new algorithms for exploiting the correlation among keywords in a filtering setting. We develop and evaluate two algorithms based on single-key and multi-key statistics and utilize two different synopses (Hash Sketches and KMV synopses) to compactly represent publishers. Our experimental evaluation using two real-life corpora with web and blog data demonstrates the filtering effectiveness of both approaches and highlights the different tradeoffs.

  • efficient search and Approximate Information filtering in a distributed peer to peer environment of digital libraries
    ACM international conference on Digital libraries, 2007
    Co-Authors: Christian Zimmer, Christos Tryfonopoulos, Gerhard Weikum
    Abstract:

    We present a new architecture for efficient search and Approximate Information filtering in a distributed Peer-to-Peer (P2P) environment of Digital Libraries. The MinervaLight search system uses P2P techniques over a structured overlay network to distribute and maintain a directory of peer statistics. Based on the same directory, the MAPS Information filtering system provides an Approximate publish/subscribe functionality by monitoring the most promising digital libraries for publishing appropriate documents regarding a continuous query. In this paper, we discuss our system architecture that combines searching and Information filtering abilities. We show the system components of MinervaLight and explain the different facets of an Approximate pub/sub system for subscriptions that is high scalable, efficient, and notifies the subscribers about the most interesting publications in the P2P network of digital libraries. We also compare both approaches in terms of common properties and differences to show an overview of search and pub/sub using the same infrastructure.

R. V. Erickson - One of the best experts on this subject based on the ideXlab platform.

James A Landay - One of the best experts on this subject based on the ideXlab platform.

  • Approximate Information flows socially based modeling of privacy in ubiquitous computing
    Ubiquitous Computing, 2002
    Co-Authors: Xiaodong Jiang, Jason Hong, James A Landay
    Abstract:

    In this paper, we propose a framework for supporting sociallycompatible privacy objectives in ubiquitous computing settings. Drawing on social science research, we have developed a key objective called the Principle of Minimum Asymmetry, which seeks to minimize the imbalance between the people about whom data is being collected, and the systems and people that collect and use that data. We have also developed Approximate Information Flow (AIF), a model describing the interaction between the various actors and personal data. AIF effectively supports varying degrees of asymmetry for ubicomp systems, suggests new privacy protection mechanisms, and provides a foundation for inspecting privacy-friendliness of ubicomp systems.

  • UbiComp - Approximate Information Flows: Socially-Based Modeling of Privacy in Ubiquitous Computing
    2002
    Co-Authors: Xiaodong Jiang, Jason Hong, James A Landay
    Abstract:

    In this paper, we propose a framework for supporting sociallycompatible privacy objectives in ubiquitous computing settings. Drawing on social science research, we have developed a key objective called the Principle of Minimum Asymmetry, which seeks to minimize the imbalance between the people about whom data is being collected, and the systems and people that collect and use that data. We have also developed Approximate Information Flow (AIF), a model describing the interaction between the various actors and personal data. AIF effectively supports varying degrees of asymmetry for ubicomp systems, suggests new privacy protection mechanisms, and provides a foundation for inspecting privacy-friendliness of ubicomp systems.

Hyunbo Cho - One of the best experts on this subject based on the ideXlab platform.