The Experts below are selected from a list of 51 Experts worldwide ranked by ideXlab platform

Zhen Zhang - One of the best experts on this subject based on the ideXlab platform.

  • toward large scale integration building a metaquerier over databases on the web
    Conference on Innovative Data Systems Research, 2005
    Co-Authors: Kevin Chen-chuan Chang, Bin He, Zhen Zhang
    Abstract:

    The Web has been rapidly “deepened” by myriad searchable databases online, where data are hidden behind query interfaces. Toward large scale integration over this “deep Web,” we have been building the MetaQuerier system– for both exploring (to find) and integrating (to query) databases on the Web. As an interim report, first, this paper proposes our goal of the MetaQuerier for Web-scale integration– With its dynamic and ad-hoc nature, such large scale integration mandates both dynamic Source Discovery and on-thefly query translation. Second, we present the system architecture and underlying technology of key subsystems in our ongoing implementation. Third, we discuss “lessons” learned to date, focusing on our efforts in system integration, for putting individual subsystems to function together. On one hand, we observe that, across subsystems, the system integration of an integration system is itself non-trivial– which presents both challenges and opportunities beyond subsystems in isolation. On the other hand, we also observe that, across subsystems, there emerge unified insights of “holistic integration”– which leverage large scale itself as a unique opportunity for information integration.

Deevakar Rogith - One of the best experts on this subject based on the ideXlab platform.

  • datamed an open Source Discovery index for finding biomedical datasets
    Journal of the American Medical Informatics Association, 2018
    Co-Authors: Xiaoling Chen, Anupama E Gururaj, Burak Ozyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Nansu Zong, Min Jiang, Deevakar Rogith
    Abstract:

    Author(s): Chen, Xiaoling; Gururaj, Anupama E; Ozyurt, Burak; Liu, Ruiling; Soysal, Ergin; Cohen, Trevor; Tiryaki, Firat; Li, Yueling; Zong, Nansu; Jiang, Min; Rogith, Deevakar; Salimi, Mandana; Kim, Hyeon-Eui; Rocca-Serra, Philippe; Gonzalez-Beltran, Alejandra; Farcas, Claudiu; Johnson, Todd; Margolis, Ron; Alter, George; Sansone, Susanna-Assunta; Fore, Ian M; Ohno-Machado, Lucila; Grethe, Jeffrey S; Xu, Hua | Abstract: ObjectiveFinding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open Source biomedical data Discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain.Materials and methodsDataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine.Results and conclusionOur manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open Source package for the biomedical community.

Jadwiga Indulska - One of the best experts on this subject based on the ideXlab platform.

  • context privacy and obfuscation supported by dynamic context Source Discovery and processing in a context management system
    Ubiquitous Intelligence and Computing, 2007
    Co-Authors: Ryan Wishart, Karen Henricksen, Jadwiga Indulska
    Abstract:

    The extensive context information collection abilities of ubiquitous computing environments represent a significant threat to user privacy. In this paper we address this threat by introducing a context information privacy mechanism. Our approach relies on context-dependent ownership definitions and context owner-specified privacy preferences to control context disclosure to third-parties. These privacy preferences enable context owners to stipulate not only to whom their context information can be disclosed and the conditions of disclosure, but also the level of detail at which the context information can be disclosed. Context information that cannot be disclosed at its existing level of detail is obfuscated to meet detail level requirements stipulated by its owner. To achieve this obfuscation of context information we introduce a new approach based on dynamic Discovery and processing of context Sources. Our new approach is demonstrated in a Context Management System in which context Source Discovery and processing is facilitated by the SensorML sensor description standard being developed by the Open Geospatial Consortium.

Kevin Chen-chuan Chang - One of the best experts on this subject based on the ideXlab platform.

  • toward large scale integration building a metaquerier over databases on the web
    Conference on Innovative Data Systems Research, 2005
    Co-Authors: Kevin Chen-chuan Chang, Bin He, Zhen Zhang
    Abstract:

    The Web has been rapidly “deepened” by myriad searchable databases online, where data are hidden behind query interfaces. Toward large scale integration over this “deep Web,” we have been building the MetaQuerier system– for both exploring (to find) and integrating (to query) databases on the Web. As an interim report, first, this paper proposes our goal of the MetaQuerier for Web-scale integration– With its dynamic and ad-hoc nature, such large scale integration mandates both dynamic Source Discovery and on-thefly query translation. Second, we present the system architecture and underlying technology of key subsystems in our ongoing implementation. Third, we discuss “lessons” learned to date, focusing on our efforts in system integration, for putting individual subsystems to function together. On one hand, we observe that, across subsystems, the system integration of an integration system is itself non-trivial– which presents both challenges and opportunities beyond subsystems in isolation. On the other hand, we also observe that, across subsystems, there emerge unified insights of “holistic integration”– which leverage large scale itself as a unique opportunity for information integration.

Li Xu - One of the best experts on this subject based on the ideXlab platform.

  • Source Discovery and schema mapping for data integration
    2003
    Co-Authors: David W Embley, Li Xu
    Abstract:

    As data explodes on the Web, there is a need to integrate data from a large number of heterogeneous information Sources. Currently, there are two main basic approaches to data integration: Global-as-View (GAV) and Local-as-View (LAV). However, both approaches have their limitations for large-scale applications. To resolve the problems, we offer a Target-based Integration Query System (TIQS) as an alternative point of view that is neither GAV nor LAV The approach uses a predefined conceptual target schema, which is specified ontologically and independently of any of the Sources, as a central, organizing concept. In this dissertation, we focus on the resolutions to three problems in TIQS: (1) automatically recognizing information Sources for the target, (2) automating Source-to-target mappings between Source and target schemas, and (3) query reformulation based on Source-to-target mappings. Experiments we have conducted show that we have been able to achieve good performance for the recognition of applicable documents as well as the generation of Source-to-target mappings. Moreover, we have proven that query reformulation in TIQS reduces to rule unfolding and the reformulated user queries extract all the query answers available from Sources with respect to the definition of TIQS for the proposed queries.