The Experts below are selected from a list of 21918 Experts worldwide ranked by ideXlab platform
Min Zhang - One of the best experts on this subject based on the ideXlab platform.
-
using Format migration and preservation Metadata to support digital preservation of scientific data
International Conference on Software Engineering, 2019Co-Authors: Jiajun Xie, Min ZhangAbstract:With the development of e-Science and data intensive scientific discovery, it needs to ensure scientific data available for the long-term, with the goal that the valuable scientific data should be discovered and re-used for downstream investigations, either alone, or in combination with newly generated data. As such, the preservation of scientific data enables that not only might experiment be reproducible and verifiable, but also new questions can be raised by other scientists to promote research and innovation. In this paper, we focus on the two main problems of digital preservation that are Format migration and preservation Metadata. Format migration includes both Format verification and object transFormation. The system architecture of Format migration and preservation Metadata is presented, mapping rules of object transFormation are analyzed, data fixity and integrity and authenticity, digital signature and so on are discussed and an example is shown in detail.
Andreas W Liehr - One of the best experts on this subject based on the ideXlab platform.
-
on the communication of scientific data the full Metadata Format
Computer Physics Communications, 2010Co-Authors: Moritz Riede, Rico Schueppel, Kristian O Sylvesterhvid, Martin Kuhne, Michael C Rottger, Klaus Zimmermann, Andreas W LiehrAbstract:In this paper, we introduce a scientific Format for text-based data files, which facilitates storing and communicating tabular data sets. The so-called Full-Metadata Format builds on the widely used INI-standard and is based on four principles: readable self-documentation, flexible structure, fail-safe compatibility, and searchability. As a consequence, all Metadata required to interpret the tabular data are stored in the same file, allowing for the automated generation of publication-ready tables and graphs and the semantic searchability of data file collections. The Full-Metadata Format is introduced on the basis of three comprehensive examples. The complete Format and syntax are given in the appendix.
H. Jeremy Bockholt - One of the best experts on this subject based on the ideXlab platform.
-
A National Human Neuroimaging Collaboratory Enabled by the Biomedical InFormatics Research Network (BIRN)
IEEE Transactions on Information Technology in Biomedicine, 2008Co-Authors: David B. Keator, S. Pieper, D. Marcus, Basak Ozyurt, D. Greve, Sunil Gadde, Randy Notestine, Sean Murphy, Jeffrey S. Grethe, H. Jeremy BockholtAbstract:The aggregation of imaging, clinical, and behavioral data from multiple independent institutions and researchers presents both a great opportunity for biomedical research as well as a formidable challenge. Many research groups have well-established data collection and analysis procedures, as well as data and Metadata Format requirements that are particular to that group. Moreover, the types of data and Metadata collected are quite diverse, including image, physiological, and behavioral data, as well as descriptions of experimental design, and preprocessing and analysis methods. Each of these types of data utilizes a variety of software tools for collection, storage, and processing. Furthermore sites are reluctant to release control over the distribution and access to the data and the tools. To address these needs, the biomedical inFormatics research network (BIRN) has developed a federated and distributed infrastructure for the storage, retrieval, analysis, and documentation of biomedical imaging data. The infrastructure consists of distributed data collections hosted on dedicated storage and computational resources located at each participating site, a federated data management system and data integration environment, an extensible markup language (XML) schema for data exchange, and analysis pipelines, designed to leverage both the distributed data management environment and the available grid computing resources.
Sprenger Julia - One of the best experts on this subject based on the ideXlab platform.
-
Tools and workflows for data & Metadata management of complex experiments : building a foundation for reproducible & collaborative analysis in the neurosciences
Forschungszentrum Jülich GmbH Zentralbibliothek Verlag, 2020Co-Authors: Sprenger JuliaAbstract:The scientific knowledge of mankind is based on the verification of hypotheses by carrying out experiments. As the construction and conduct of an experiment becomes increasingly complex more and more scientists are involved in a single project. In order to make the generated data easily accessible to all scientists and, at best, to the entire scientific community, it is essential to comprehensively document the circumstances of the data generation, as these contain essential inFormation for later analysis and interpretation. In this thesis, I present two complex neuroscience projects and the strategies, tools, and concepts that were used to comprehensively track, process, organize, and prepare the collected data for joint analysis. First, I describe the older of the two experiments and explain in detail the generation of data and Metadata and the pipeline used for aggregating Metadata. A hierarchical approach based on the open source software odMLfor Metadata organization was implemented to capture the complex meta inFormation of this project. I evaluate the design concepts and tools used and derive a general catalogue of requirements for scientific collaboration in complex projects. Also, I identify issues and requirements that were not yet addressed by this pipeline. There were, in particular, the difficulties in i) entering manual Metadata and structuring the Metadata collection, ii) combining Metadata with the actual data, and iii) setting up the pipeline in a modular generic and transparent manner. Guided by this analysis, I describe concept and tool implementations to address these identified issues. I developed a complementary tool (odMLtables) to i) facilitate the capture of Metadata in a structured way and to ii) convert these easily into the hierarchical, standardized Metadata Format odML. odMLtables provides an interface between the easy-to-read tabular Metadata representation in the Formats commonly used in lab-oratory environments (csv/xls) and the hierarchically organized odML Format based on xml, which is designed for a comprehensive collection of complex Metadata records in an easily machine-readable manner. Supplementing the coordinated capture of Metadata, I contributed to and shaped the Neo toolbox for the standardized representation of electrophysiological data. This toolbox is a key component for electrophysiological data analysis as it integrates different proprietary and non-proprietary file Formats and serves as a bridge between different file Formats. I emphasize new features that simplify the process of data and Metadata handling in the data acquisition workflow. I introduce the concept of workflow management into the field of scientific data pro-cessing, based on the common Python-based snake make package. For the second, more recent electrophysiological experiment, I designed and implemented the workflow for capturing and packaging Metadata and data in a comprehensive form. Here I used the generic neuroscience inFormation exchange Format (Nix) for the user-friendly packaging of data sets including data and Metadata in combined form. Finally, I evaluate the improved workflow against the requirements of collaborative scientific work in complex projects. I establish general guidelines for conducting such experiments and workflows in a scientific environment. In conclusion, I present the next development steps for the presented workflow and potential avenues for deploying this prototype as a production prototype to a wider scientific community
-
Tools and Workflows for Data & Metadata Management of Complex Experiments - Building a Foundation for Reproducible & Collaborative Analysis in the Neurosciences
Forschungszentrum Jülich GmbH Zentralbibliothek Verlag, 2020Co-Authors: Sprenger JuliaAbstract:The scientific knowledge of mankind is based on the verification of hypotheses by carrying out experiments. As the construction and conduct of an experiment becomes increasingly complex more and more scientists are involved in a single project. In order to make the generated data easily accessible to all scientists and, at best, to the entire scientific community, it is essential to comprehensively document the circumstances of the data generation, as these contain essential inFormation for later analysis and interpretation. In this thesis, I present two complex neuroscience projects and the strategies, tools, and concepts that were used to comprehensively track, process, organize, and prepare the collected data for joint analysis. First, I describe the older of the two experiments and explain in detail the generation of data and Metadata and the pipeline used for aggregating Metadata. A hierarchical approach based on the open source software $\textit{odML}$ for Metadata organization was implemented to capture the complex meta inFormation of this project. I evaluate the design concepts and tools used and derive a general catalogue of requirements for scientific collaboration in complex projects. Also, I identify issues and requirements that were not yet addressed by this pipeline. There were, in particular, the difficulties in i) entering manual Metadata and structuring the Metadata collection,ii) combining Metadata with the actual data, and iii) setting up the pipeline in a modular generic and transparent manner. Guided by this analysis, I describe concept and tool implementations to address these identified issues. I developed a complementary tool ($\textit{odMLtables}$) to i) facilitate the capture of Metadata in a structured way and to ii) convert these easily into the hierarchical, standardized Metadata Format $\textit{odML. odMLtables}$ provides an interface between the easy-to-read tabular Metadata representation in the Formats commonly used in laboratory environments (csv/xls) and the hierarchically organized $\textit{odML}$ Format based on xml, which is designed for a comprehensive collection of complex Metadata records in an easily machine-readable manner. Supplementing the coordinated capture of Metadata, I contributed to and shaped the $\textit{Neo}$ toolbox for the standardized representation of electrophysiological data. This toolbox is a key component for electrophysiological data analysis as it integrates different proprietary and non-proprietary file Formats and serves as a bridge between different file Formats. I emphasize new features that simplify the process of data and Metadata handling in the data acquisition workflow. I introduce the concept of workflow management into the field of scientific data processing, based on the common Python-based snakemake package. For the second, more recent electrophysiological experiment, I designed and implemented the workflow for capturing and packaging Metadata and data in a comprehensive form. Here I used the generic neuroscience inFormation exchange Format ($\textit{Nix}$) for the user-friendly packaging of data sets including data and Metadata in combined form. Finally, I evaluate the improved workflow against the requirements of collaborative scientific work in complex projects. I establish general guidelines for conducting such experiments and workflows in a scientific environment. In conclusion, I present the next development steps for the presented workflow and potential avenues for deploying this prototype as a production prototype to a wider scientific community
Jonas S Almeida - One of the best experts on this subject based on the ideXlab platform.
-
rppaml rims a Metadata Format and an inFormation management system for reverse phase protein arrays
BMC Bioinformatics, 2008Co-Authors: Romesh Stanislaus, Mark S Carey, Helena F Deus, Kevin R Coombes, Bryan T Hennessy, Gordon B Mills, Jonas S AlmeidaAbstract:Reverse Phase Protein Arrays (RPPA) are convenient assay platforms to investigate the presence of biomarkers in tissue lysates. As with other high-throughput technologies, substantial amounts of analytical data are generated. Over 1000 samples may be printed on a single nitrocellulose slide. Up to 100 different proteins may be assessed using immunoperoxidase or immunoflorescence techniques in order to determine relative amounts of protein expression in the samples of interest. In this report an RPPA InFormation Management System (RIMS) is described and made available with open source software. In order to implement the proposed system, we propose a Metadata Format known as reverse phase protein array markup language (RPPAML). RPPAML would enable researchers to describe, document and disseminate RPPA data. The complexity of the data structure needed to describe the results and the graphic tools necessary to visualize them require a software deployment distributed between a client and a server application. This was achieved without sacrificing interoperability between individual deployments through the use of an open source semantic database, S3DB. This data service backbone is available to multiple client side applications that can also access other server side deployments. The RIMS platform was designed to interoperate with other data analysis and data visualization tools such as Cytoscape. The proposed RPPAML data Format hopes to standardize RPPA data. Standardization of data would result in diverse client applications being able to operate on the same set of data. Additionally, having data in a standard Format would enable data dissemination and data analysis.