The Experts below are selected from a list of 60 Experts worldwide ranked by ideXlab platform

Camille Mathieu - One of the best experts on this subject based on the ideXlab platform.

  • practical application of the dublin core standard for enterprise Metadata management
    Association for Information Science and Technology, 2017
    Co-Authors: Camille Mathieu
    Abstract:

    EDITOR'S SUMMARY Large organizations relying heavily on knowledge work require effective capture and reuse of information, enabled through consistent use of standardized enterprise content Metadata. The Jet Propulsion Laboratory (JPL) has undertaken a standardization effort, building an internal content schema based on established Metadata Field standards that are content- and application-agnostic but locally customizable for application to a broad variety of repositories. The JPL adopted the Dublin Core standard, with its Simple and Qualified properties as well as further refined Custom sub-properties. The JPL Resource Schema serves as an enterprise-wide Metadata standard, while specific application profiles state the available Fields and Field labels for each repository or content management system. The schema's terms are drawn from two distinct but semantically related vocabularies and linked by an intermediary registry tying granular listings for specific applications to enterprise-level terms. The registry mappings permit the use of both local Metadata and higher level or external systems. The effort has demonstrated the importance of consistent application of both granular and general Metadata for information capture and revealed important lessons about adopting the Dublin Core standard in a large enterprise setting.

Gareth J. F. Jones - One of the best experts on this subject based on the ideXlab platform.

  • Utilisation of Metadata Fields and query expansion in cross-lingual search of user-generated internet video
    Journal of Artificial Intelligence Research, 2016
    Co-Authors: Ahmad Khwileh, Debasis Ganguly, Gareth J. F. Jones
    Abstract:

    Recent years have seen significant efforts in the area of Cross Language Information Retrieval (CLIR) for text retrieval. This work initially focused on formally published content, but more recently research has begun to concentrate on CLIR for informal social media content. However, despite the current expansion in online multimedia archives, there has been little work on CLIR for this content. While there has been some limited work on Cross-Language Video Retrieval (CLVR) for professional videos, such as documentaries or TV news broadcasts, there has to date, been no significant investigation of CLVR for the rapidly growing archives of informal user generated (UGC) content. Key differences between such UGC and professionally produced content are the nature and structure of the textual UGC Metadata associated with it, as well as the form and quality of the content itself. In this setting, retrieval effectiveness may not only suffer from translation errors common to all CLIR tasks, but also recognition errors associated with the automatic speech recognition (ASR) systems used to transcribe the spoken content of the video and with the informality and inconsistency of the associated user-created Metadata for each video. This work proposes and evaluates techniques to improve CLIR effectiveness of such noisy UGC content. Our experimental investigation shows that different sources of evidence, e.g. the content from different Fields of the structured Metadata, significantly affect CLIR effectiveness. Results from our experiments also show that each Metadata Field has a varying robustness to query expansion (QE) and hence can have a negative impact on the CLIR effectiveness. Our work proposes a novel adaptive QE technique that predicts the most reliable source for expansion and shows how this technique can be effective for improving CLIR effectiveness for UGC content.

Ahmad Khwileh - One of the best experts on this subject based on the ideXlab platform.

  • Utilisation of Metadata Fields and query expansion in cross-lingual search of user-generated internet video
    Journal of Artificial Intelligence Research, 2016
    Co-Authors: Ahmad Khwileh, Debasis Ganguly, Gareth J. F. Jones
    Abstract:

    Recent years have seen significant efforts in the area of Cross Language Information Retrieval (CLIR) for text retrieval. This work initially focused on formally published content, but more recently research has begun to concentrate on CLIR for informal social media content. However, despite the current expansion in online multimedia archives, there has been little work on CLIR for this content. While there has been some limited work on Cross-Language Video Retrieval (CLVR) for professional videos, such as documentaries or TV news broadcasts, there has to date, been no significant investigation of CLVR for the rapidly growing archives of informal user generated (UGC) content. Key differences between such UGC and professionally produced content are the nature and structure of the textual UGC Metadata associated with it, as well as the form and quality of the content itself. In this setting, retrieval effectiveness may not only suffer from translation errors common to all CLIR tasks, but also recognition errors associated with the automatic speech recognition (ASR) systems used to transcribe the spoken content of the video and with the informality and inconsistency of the associated user-created Metadata for each video. This work proposes and evaluates techniques to improve CLIR effectiveness of such noisy UGC content. Our experimental investigation shows that different sources of evidence, e.g. the content from different Fields of the structured Metadata, significantly affect CLIR effectiveness. Results from our experiments also show that each Metadata Field has a varying robustness to query expansion (QE) and hence can have a negative impact on the CLIR effectiveness. Our work proposes a novel adaptive QE technique that predicts the most reliable source for expansion and shows how this technique can be effective for improving CLIR effectiveness for UGC content.

Mark A. Musen - One of the best experts on this subject based on the ideXlab platform.

  • The variable quality of Metadata about biological samples used in biomedical experiments.
    Scientific data, 2019
    Co-Authors: Rafael S. Gonçalves, Mark A. Musen
    Abstract:

    We present an analytical study of the quality of Metadata about samples used in biomedical experiments. The Metadata under analysis are stored in two well-known databases: BioSample—a repository managed by the National Center for Biotechnology Information (NCBI), and BioSamples—a repository managed by the European Bioinformatics Institute (EBI). We tested whether 11.4 M sample Metadata records in the two repositories are populated with values that fulfill the stated requirements for such values. Our study revealed multiple anomalies in the Metadata. Most Metadata Field names and their values are not standardized or controlled. Even simple binary or numeric Fields are often populated with inadequate values of different data types. By clustering Metadata Field names, we discovered there are often many distinct ways to represent the same aspect of a sample. Overall, the Metadata we analyzed reveal that there is a lack of principled mechanisms to enforce and validate Metadata requirements. The significant aberrancies that we found in the Metadata are likely to impede search and secondary use of the associated datasets.

  • Metadata in the BioSample Online Repository are Impaired by Numerous Anomalies
    arXiv: Databases, 2017
    Co-Authors: Rafael S. Gonçalves, Martin J. O'connor, Marcos Martínez-romero, John Graybeal, Mark A. Musen
    Abstract:

    The Metadata about scientific experiments are crucial for finding, reproducing, and reusing the data that the Metadata describe. We present a study of the quality of the Metadata stored in BioSample--a repository of Metadata about samples used in biomedical experiments managed by the U.S. National Center for Biomedical Technology Information (NCBI). We tested whether 6.6 million BioSample Metadata records are populated with values that fulfill the stated requirements for such values. Our study revealed multiple anomalies in the analyzed Metadata. The BioSample Metadata Field names and their values are not standardized or controlled--15% of the Metadata Fields use Field names not specified in the BioSample data dictionary. Only 9 out of 452 BioSample-specified Fields ordinarily require ontology terms as values, and the quality of these controlled Fields is better than that of uncontrolled ones, as even simple binary or numeric Fields are often populated with inadequate values of different data types (e.g., only 27% of Boolean values are valid). Overall, the Metadata in BioSample reveal that there is a lack of principled mechanisms to enforce and validate Metadata requirements. The aberrancies in the Metadata are likely to impede search and secondary use of the associated datasets.

Debasis Ganguly - One of the best experts on this subject based on the ideXlab platform.

  • Utilisation of Metadata Fields and query expansion in cross-lingual search of user-generated internet video
    Journal of Artificial Intelligence Research, 2016
    Co-Authors: Ahmad Khwileh, Debasis Ganguly, Gareth J. F. Jones
    Abstract:

    Recent years have seen significant efforts in the area of Cross Language Information Retrieval (CLIR) for text retrieval. This work initially focused on formally published content, but more recently research has begun to concentrate on CLIR for informal social media content. However, despite the current expansion in online multimedia archives, there has been little work on CLIR for this content. While there has been some limited work on Cross-Language Video Retrieval (CLVR) for professional videos, such as documentaries or TV news broadcasts, there has to date, been no significant investigation of CLVR for the rapidly growing archives of informal user generated (UGC) content. Key differences between such UGC and professionally produced content are the nature and structure of the textual UGC Metadata associated with it, as well as the form and quality of the content itself. In this setting, retrieval effectiveness may not only suffer from translation errors common to all CLIR tasks, but also recognition errors associated with the automatic speech recognition (ASR) systems used to transcribe the spoken content of the video and with the informality and inconsistency of the associated user-created Metadata for each video. This work proposes and evaluates techniques to improve CLIR effectiveness of such noisy UGC content. Our experimental investigation shows that different sources of evidence, e.g. the content from different Fields of the structured Metadata, significantly affect CLIR effectiveness. Results from our experiments also show that each Metadata Field has a varying robustness to query expansion (QE) and hence can have a negative impact on the CLIR effectiveness. Our work proposes a novel adaptive QE technique that predicts the most reliable source for expansion and shows how this technique can be effective for improving CLIR effectiveness for UGC content.