The Experts below are selected from a list of 78 Experts worldwide ranked by ideXlab platform
Nagiza F Samatova - One of the best experts on this subject based on the ideXlab platform.
-
a network fusion guided dashboard interface for task Centric Document curation
Intelligent User Interfaces, 2017Co-Authors: Paul Jones, Shivani Sharma, Changsung Moon, Nagiza F SamatovaAbstract:Knowledge workers are being exposed to more information than ever before, as well as having to work in multi-tasking and collaborative environments. There is an increasing need for interfaces and algorithms to help automatically keep track of Documents that are associated with both individual and team tasks. Previous approaches to the problem of automatically applying task labels to Documents have been limited to small feature spaces or have not taken into account multi-user environments. Many different clues to potential task associations are available through user, task and Document similarity metrics, as well as through temporal patterns in individual and team workflows. We present a network-fusion algorithm for automatic task-Centric Document curation, and show how this can guide a recent-work dashboard interface, which organizes user's Documents and gathers feedback from them. Our approach efficiently computes representations of users, tasks and Documents in a common vector space, and can easily take into account many different types of associations through the creation of edges in a multi-layer graph. We have demonstrated the effectiveness of this approach using labelled Document corpora from three empirical studies with students and intelligence analysts. We have also shown how to leverage relationships between different entity types to increase classification accuracy by up to 20% over a simpler baseline, and with as little as 10% labelled data.
-
IUI - A Network-Fusion Guided Dashboard Interface for Task-Centric Document Curation
Proceedings of the 22nd International Conference on Intelligent User Interfaces, 2017Co-Authors: Paul Jones, Shivani Sharma, Changsung Moon, Nagiza F SamatovaAbstract:Knowledge workers are being exposed to more information than ever before, as well as having to work in multi-tasking and collaborative environments. There is an increasing need for interfaces and algorithms to help automatically keep track of Documents that are associated with both individual and team tasks. Previous approaches to the problem of automatically applying task labels to Documents have been limited to small feature spaces or have not taken into account multi-user environments. Many different clues to potential task associations are available through user, task and Document similarity metrics, as well as through temporal patterns in individual and team workflows. We present a network-fusion algorithm for automatic task-Centric Document curation, and show how this can guide a recent-work dashboard interface, which organizes user's Documents and gathers feedback from them. Our approach efficiently computes representations of users, tasks and Documents in a common vector space, and can easily take into account many different types of associations through the creation of edges in a multi-layer graph. We have demonstrated the effectiveness of this approach using labelled Document corpora from three empirical studies with students and intelligence analysts. We have also shown how to leverage relationships between different entity types to increase classification accuracy by up to 20% over a simpler baseline, and with as little as 10% labelled data.
Elena Demidova - One of the best experts on this subject based on the ideXlab platform.
-
Towards extracting event-Centric collections from Web archives
International Journal on Digital Libraries, 2020Co-Authors: Gerhard Gossen, Thomas Risse, Elena DemidovaAbstract:Web archives constitute an increasingly important source of information for computer scientists, humanities researchers and journalists interested in studying past events. However, currently there are no access methods that help Web archive users to efficiently access event-Centric information in large-scale archives that go beyond the retrieval of individual disconnected Documents. In this article, we tackle the novel problem of extracting interlinked event-Centric Document collections from large-scale Web archives to facilitate an efficient and intuitive access to information regarding past events. We address this problem by: (1) facilitating users to define event-Centric Document collections in an intuitive way through a Collection Specification; (2) development of a specialised extraction method that adapts focused crawling techniques to the Web archive settings; and (3) definition of a function to judge the relevance of the archived Documents with respect to the Collection Specification taking into account the topical and temporal relevance of the Documents. Our extended experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables efficient extraction of event-Centric collections for different event types.
-
TPDL - Extracting Event-Centric Document Collections from Large-Scale Web Archives
Research and Advanced Technology for Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.
-
Extracting Event-Centric Document Collections from Large-Scale Web Archives
arXiv: Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.
Gerhard Gossen - One of the best experts on this subject based on the ideXlab platform.
-
Towards extracting event-Centric collections from Web archives
International Journal on Digital Libraries, 2020Co-Authors: Gerhard Gossen, Thomas Risse, Elena DemidovaAbstract:Web archives constitute an increasingly important source of information for computer scientists, humanities researchers and journalists interested in studying past events. However, currently there are no access methods that help Web archive users to efficiently access event-Centric information in large-scale archives that go beyond the retrieval of individual disconnected Documents. In this article, we tackle the novel problem of extracting interlinked event-Centric Document collections from large-scale Web archives to facilitate an efficient and intuitive access to information regarding past events. We address this problem by: (1) facilitating users to define event-Centric Document collections in an intuitive way through a Collection Specification; (2) development of a specialised extraction method that adapts focused crawling techniques to the Web archive settings; and (3) definition of a function to judge the relevance of the archived Documents with respect to the Collection Specification taking into account the topical and temporal relevance of the Documents. Our extended experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables efficient extraction of event-Centric collections for different event types.
-
TPDL - Extracting Event-Centric Document Collections from Large-Scale Web Archives
Research and Advanced Technology for Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.
-
Extracting Event-Centric Document Collections from Large-Scale Web Archives
arXiv: Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.
Paul Jones - One of the best experts on this subject based on the ideXlab platform.
-
a network fusion guided dashboard interface for task Centric Document curation
Intelligent User Interfaces, 2017Co-Authors: Paul Jones, Shivani Sharma, Changsung Moon, Nagiza F SamatovaAbstract:Knowledge workers are being exposed to more information than ever before, as well as having to work in multi-tasking and collaborative environments. There is an increasing need for interfaces and algorithms to help automatically keep track of Documents that are associated with both individual and team tasks. Previous approaches to the problem of automatically applying task labels to Documents have been limited to small feature spaces or have not taken into account multi-user environments. Many different clues to potential task associations are available through user, task and Document similarity metrics, as well as through temporal patterns in individual and team workflows. We present a network-fusion algorithm for automatic task-Centric Document curation, and show how this can guide a recent-work dashboard interface, which organizes user's Documents and gathers feedback from them. Our approach efficiently computes representations of users, tasks and Documents in a common vector space, and can easily take into account many different types of associations through the creation of edges in a multi-layer graph. We have demonstrated the effectiveness of this approach using labelled Document corpora from three empirical studies with students and intelligence analysts. We have also shown how to leverage relationships between different entity types to increase classification accuracy by up to 20% over a simpler baseline, and with as little as 10% labelled data.
-
IUI - A Network-Fusion Guided Dashboard Interface for Task-Centric Document Curation
Proceedings of the 22nd International Conference on Intelligent User Interfaces, 2017Co-Authors: Paul Jones, Shivani Sharma, Changsung Moon, Nagiza F SamatovaAbstract:Knowledge workers are being exposed to more information than ever before, as well as having to work in multi-tasking and collaborative environments. There is an increasing need for interfaces and algorithms to help automatically keep track of Documents that are associated with both individual and team tasks. Previous approaches to the problem of automatically applying task labels to Documents have been limited to small feature spaces or have not taken into account multi-user environments. Many different clues to potential task associations are available through user, task and Document similarity metrics, as well as through temporal patterns in individual and team workflows. We present a network-fusion algorithm for automatic task-Centric Document curation, and show how this can guide a recent-work dashboard interface, which organizes user's Documents and gathers feedback from them. Our approach efficiently computes representations of users, tasks and Documents in a common vector space, and can easily take into account many different types of associations through the creation of edges in a multi-layer graph. We have demonstrated the effectiveness of this approach using labelled Document corpora from three empirical studies with students and intelligence analysts. We have also shown how to leverage relationships between different entity types to increase classification accuracy by up to 20% over a simpler baseline, and with as little as 10% labelled data.
Thomas Risse - One of the best experts on this subject based on the ideXlab platform.
-
Towards extracting event-Centric collections from Web archives
International Journal on Digital Libraries, 2020Co-Authors: Gerhard Gossen, Thomas Risse, Elena DemidovaAbstract:Web archives constitute an increasingly important source of information for computer scientists, humanities researchers and journalists interested in studying past events. However, currently there are no access methods that help Web archive users to efficiently access event-Centric information in large-scale archives that go beyond the retrieval of individual disconnected Documents. In this article, we tackle the novel problem of extracting interlinked event-Centric Document collections from large-scale Web archives to facilitate an efficient and intuitive access to information regarding past events. We address this problem by: (1) facilitating users to define event-Centric Document collections in an intuitive way through a Collection Specification; (2) development of a specialised extraction method that adapts focused crawling techniques to the Web archive settings; and (3) definition of a function to judge the relevance of the archived Documents with respect to the Collection Specification taking into account the topical and temporal relevance of the Documents. Our extended experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables efficient extraction of event-Centric collections for different event types.
-
TPDL - Extracting Event-Centric Document Collections from Large-Scale Web Archives
Research and Advanced Technology for Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.
-
Extracting Event-Centric Document Collections from Large-Scale Web Archives
arXiv: Digital Libraries, 2017Co-Authors: Gerhard Gossen, Elena Demidova, Thomas RisseAbstract:Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in the social sciences, the historical sciences and journalists interested in studying past events. However, there are currently no access methods that help users to efficiently access information, in particular about specific events, beyond the retrieval of individual disconnected Documents. Therefore we propose a novel method to extract event-Centric Document collections from large scale Web archives. This method relies on a specialized focused extraction algorithm. Our experiments on the German Web archive (covering a time period of 19 years) demonstrate that our method enables the extraction of event-Centric collections for different event types.