The Experts below are selected from a list of 2505 Experts worldwide ranked by ideXlab platform

Auguste Genovesio - One of the best experts on this subject based on the ideXlab platform.

  • ALFA: annotation landscape for aligned reads
    BMC Genomics, 2019
    Co-Authors: Mathieu Bahin, Benoit F. Noël, Valentine Murigneux, Charles Bernard, Leila Bastianelli, Hervé Hir, Alice Lebreton, Auguste Genovesio
    Abstract:

    Background The last 10 years have seen the rise of countless functional genomics studies based on Next-Generation Sequencing (NGS). In the vast majority of cases, whatever the species, whatever the experiment, the two first steps of data analysis consist of a quality control of the raw reads followed by a mapping of those reads to a reference genome/transcriptome. Subsequent steps then depend on the type of study that is being made. While some tools have been proposed for investigating data quality after the mapping step, there is no commonly adopted framework that would be easy to use and broadly applicable to any NGS data type. Results We present ALFA, a simple but universal tool that can be used after the mapping step on any kind of NGS experiment data for any organism with available genomic annotations. In a Single Command Line, ALFA can compute and display distribution of reads by categories (exon, intron, UTR, etc.) and biotypes (protein coding, miRNA, etc.) for a given aligned dataset with nucleotide precision. We present applications of ALFA to Ribo-Seq and RNA-Seq on Homo sapiens , CLIP-Seq on Mus musculus , RNA-Seq on Saccharomyces cerevisiae , Bisulfite sequencing on Arabidopsis thaliana and ChIP-Seq on Caenorhabditis elegans . Conclusions We show that ALFA provides a powerful and broadly applicable approach for post mapping quality control and to produce a global overview using common or dedicated annotations. It is made available to the community as an easy to install Command Line tool and from the Galaxy Tool Shed.

Niyaz Ahmed - One of the best experts on this subject based on the ideXlab platform.

  • contig layout authenticator cla a combinatorial approach to ordering and scaffolding of bacterial contigs for comparative genomics and molecular epidemiology
    PLOS ONE, 2016
    Co-Authors: Sabiha Shaik, Narender Kumar, Aditya Kumar Lankapalli, Sumeet K Tiwari, Ramani Baddam, Niyaz Ahmed
    Abstract:

    A wide variety of genome sequencing platforms have emerged in the recent past. High-throughput platforms like Illumina and 454 are essentially adaptations of the shotgun approach generating millions of fragmented Single or paired sequencing reads. To reconstruct whole genomes, the reads have to be assembled into contigs, which often require further downstream processing. The contigs can be directly ordered according to a reference, scaffolded based on paired read information, or assembled using a combination of the two approaches. While the reference-based approach appears to mask strain-specific information, scaffolding based on paired-end information suffers when repetitive elements longer than the size of the sequencing reads are present in the genome. Sequencing technologies that produce long reads can solve the problems associated with repetitive elements but are not necessarily easily available to researchers. The most common high-throughput technology currently used is the Illumina short read platform. To improve upon the shortcomings associated with the construction of draft genomes with Illumina paired-end sequencing, we developed Contig-Layout-Authenticator (CLA). The CLA pipeLine can scaffold reference-sorted contigs based on paired reads, resulting in better assembled genomes. Moreover, CLA also hints at probable misassemblies and contaminations, for the users to cross-check before constructing the consensus draft. The CLA pipeLine was designed and trained extensively on various bacterial genome datasets for the ordering and scaffolding of large repetitive contigs. The tool has been validated and compared favorably with other widely-used scaffolding and ordering tools using both simulated and real sequence datasets. CLA is a user friendly tool that requires a Single Command Line input to generate ordered scaffolds.

Mathieu Bahin - One of the best experts on this subject based on the ideXlab platform.

  • ALFA: annotation landscape for aligned reads
    BMC Genomics, 2019
    Co-Authors: Mathieu Bahin, Benoit F. Noël, Valentine Murigneux, Charles Bernard, Leila Bastianelli, Hervé Hir, Alice Lebreton, Auguste Genovesio
    Abstract:

    Background The last 10 years have seen the rise of countless functional genomics studies based on Next-Generation Sequencing (NGS). In the vast majority of cases, whatever the species, whatever the experiment, the two first steps of data analysis consist of a quality control of the raw reads followed by a mapping of those reads to a reference genome/transcriptome. Subsequent steps then depend on the type of study that is being made. While some tools have been proposed for investigating data quality after the mapping step, there is no commonly adopted framework that would be easy to use and broadly applicable to any NGS data type. Results We present ALFA, a simple but universal tool that can be used after the mapping step on any kind of NGS experiment data for any organism with available genomic annotations. In a Single Command Line, ALFA can compute and display distribution of reads by categories (exon, intron, UTR, etc.) and biotypes (protein coding, miRNA, etc.) for a given aligned dataset with nucleotide precision. We present applications of ALFA to Ribo-Seq and RNA-Seq on Homo sapiens , CLIP-Seq on Mus musculus , RNA-Seq on Saccharomyces cerevisiae , Bisulfite sequencing on Arabidopsis thaliana and ChIP-Seq on Caenorhabditis elegans . Conclusions We show that ALFA provides a powerful and broadly applicable approach for post mapping quality control and to produce a global overview using common or dedicated annotations. It is made available to the community as an easy to install Command Line tool and from the Galaxy Tool Shed.

Christoph Dieterich - One of the best experts on this subject based on the ideXlab platform.

  • flexbar flexible barcode and adapter processing for next generation sequencing platforms
    Biology, 2012
    Co-Authors: Matthias Dodt, Johannes T Roehr, Rina Ahmed, Christoph Dieterich
    Abstract:

    Quantitative and systems biology approaches benefit from the unprecedented depth of next-generation sequencing. A typical experiment yields millions of short reads, which oftentimes carry particular sequence tags. These tags may be: (a) specific to the sequencing platform and library construction method (e.g., adapter sequences); (b) have been introduced by experimental design (e.g., sample barcodes); or (c) constitute some biological signal (e.g., splice leader sequences in nematodes). Our software FLEXBAR enables accurate recognition, sorting and trimming of sequence tags with maximal flexibility, based on exact overlap sequence alignment. The software supports data formats from all current sequencing platforms, including color-space reads. FLEXBAR maintains read pairings and processes separate barcode reads on demand. Our software facilitates the fine-grained adjustment of sequence tag detection parameters and search regions. FLEXBAR is a multi-threaded software and combines speed with precision. Even complex read processing scenarios might be executed with a Single Command Line call. We demonstrate the utility of the software in terms of read mapping applications, library demultiplexing and splice leader detection. FLEXBAR and additional information is available for academic use from the website: http://sourceforge.net/projects/flexbar/.

Sabiha Shaik - One of the best experts on this subject based on the ideXlab platform.

  • contig layout authenticator cla a combinatorial approach to ordering and scaffolding of bacterial contigs for comparative genomics and molecular epidemiology
    PLOS ONE, 2016
    Co-Authors: Sabiha Shaik, Narender Kumar, Aditya Kumar Lankapalli, Sumeet K Tiwari, Ramani Baddam, Niyaz Ahmed
    Abstract:

    A wide variety of genome sequencing platforms have emerged in the recent past. High-throughput platforms like Illumina and 454 are essentially adaptations of the shotgun approach generating millions of fragmented Single or paired sequencing reads. To reconstruct whole genomes, the reads have to be assembled into contigs, which often require further downstream processing. The contigs can be directly ordered according to a reference, scaffolded based on paired read information, or assembled using a combination of the two approaches. While the reference-based approach appears to mask strain-specific information, scaffolding based on paired-end information suffers when repetitive elements longer than the size of the sequencing reads are present in the genome. Sequencing technologies that produce long reads can solve the problems associated with repetitive elements but are not necessarily easily available to researchers. The most common high-throughput technology currently used is the Illumina short read platform. To improve upon the shortcomings associated with the construction of draft genomes with Illumina paired-end sequencing, we developed Contig-Layout-Authenticator (CLA). The CLA pipeLine can scaffold reference-sorted contigs based on paired reads, resulting in better assembled genomes. Moreover, CLA also hints at probable misassemblies and contaminations, for the users to cross-check before constructing the consensus draft. The CLA pipeLine was designed and trained extensively on various bacterial genome datasets for the ordering and scaffolding of large repetitive contigs. The tool has been validated and compared favorably with other widely-used scaffolding and ordering tools using both simulated and real sequence datasets. CLA is a user friendly tool that requires a Single Command Line input to generate ordered scaffolds.