The Experts below are selected from a list of 3573 Experts worldwide ranked by ideXlab platform

Muhammad Arslan - One of the best experts on this subject based on the ideXlab platform.

  • rna seq analysis of soft rush juncus effusus transcriptome sequencing de novo assembly annotation and polymorphism identification
    BMC Genomics, 2019
    Co-Authors: Muhammad Arslan, Upendra K Devisetty, Martin Porsch, Ivo Grose, Jochen A Muller, Stefan G Michalski
    Abstract:

    Juncus effusus L. (family: Juncaceae; order: Poales) is a helophytic rush growing in temperate damp or wet terrestrial habitats and is of almost cosmopolitan distribution. The species has been studied intensively with respect to its interaction with co-occurring plants as well as microbes being involved in major biogeochemical cycles. J. effusus has biotechnological value as component of Constructed Wetlands where the plant has been employed in phytoremediation of contaminated water. Its genome has not been sequenced. In this study we carried out functional annotation and polymorphism analysis of de novo assembled RNA-Seq data from 18 genotypes using 249 million paired-end Illumina HiSeq reads and 2.8 million 454 Titanium reads. The assembly comprised 158,591 contigs with a mean contig length of 780 bp. The assembly was annotated using the dammit! annotation pipeline, which queries the databases OrthoDB, Pfam-A, Rfam, and runs BUSCO (Benchmarking Single-Copy Ortholog genes). In total, 111,567 contigs (70.3%) were annotated with functional descriptions, assigned gene ontology terms, and conserved protein domains, which resulted in 30,932 non-redundant gene sequences. Results of BUSCO and KEGG pathway analyses were similar for J. effusus as for the well-studied members of the Poales, Oryza sativa and Sorghum bicolor. A total of 566,433 polymorphisms were identified in transcribed regions with an average frequency of 1 polymorphism in every 171 bases. The transcriptome assembly was of high quality and genome coverage was sufficient for global analyses. This annotated knowledge resource can be utilized for future gene expression analysis, genomic feature comparisons, genotyping, primer design, and functional genomics in J. effusus.

Stefan G Michalski - One of the best experts on this subject based on the ideXlab platform.

  • rna seq analysis of soft rush juncus effusus transcriptome sequencing de novo assembly annotation and polymorphism identification
    BMC Genomics, 2019
    Co-Authors: Muhammad Arslan, Upendra K Devisetty, Martin Porsch, Ivo Grose, Jochen A Muller, Stefan G Michalski
    Abstract:

    Juncus effusus L. (family: Juncaceae; order: Poales) is a helophytic rush growing in temperate damp or wet terrestrial habitats and is of almost cosmopolitan distribution. The species has been studied intensively with respect to its interaction with co-occurring plants as well as microbes being involved in major biogeochemical cycles. J. effusus has biotechnological value as component of Constructed Wetlands where the plant has been employed in phytoremediation of contaminated water. Its genome has not been sequenced. In this study we carried out functional annotation and polymorphism analysis of de novo assembled RNA-Seq data from 18 genotypes using 249 million paired-end Illumina HiSeq reads and 2.8 million 454 Titanium reads. The assembly comprised 158,591 contigs with a mean contig length of 780 bp. The assembly was annotated using the dammit! annotation pipeline, which queries the databases OrthoDB, Pfam-A, Rfam, and runs BUSCO (Benchmarking Single-Copy Ortholog genes). In total, 111,567 contigs (70.3%) were annotated with functional descriptions, assigned gene ontology terms, and conserved protein domains, which resulted in 30,932 non-redundant gene sequences. Results of BUSCO and KEGG pathway analyses were similar for J. effusus as for the well-studied members of the Poales, Oryza sativa and Sorghum bicolor. A total of 566,433 polymorphisms were identified in transcribed regions with an average frequency of 1 polymorphism in every 171 bases. The transcriptome assembly was of high quality and genome coverage was sufficient for global analyses. This annotated knowledge resource can be utilized for future gene expression analysis, genomic feature comparisons, genotyping, primer design, and functional genomics in J. effusus.

Paul P Gardner - One of the best experts on this subject based on the ideXlab platform.

  • Current Protocols in Bioinformatics - Studying RNA Homology and Conservation with Infernal: From Single Sequences to RNA Families
    Current Protocols in Bioinformatics, 2016
    Co-Authors: Lars Barquist, Sarah W. Burge, Paul P Gardner
    Abstract:

    Emerging high-throughput technologies have led to a deluge of putative non-coding RNA (ncRNA) sequences identified in a wide variety of organisms. Systematic characterization of these transcripts will be a tremendous challenge. Homology detection is critical to making maximal use of functional information gathered about ncRNAs: identifying homologous sequence allows us to transfer information gathered in one organism to another quickly and with a high degree of confidence. ncRNA presents a challenge for homology detection, as the primary sequence is often poorly conserved and de novo secondary structure prediction and search remain difficult. This unit introduces methods developed by the Rfam database for identifying “families” of homologous ncRNAs starting from single “seed” sequences, using manually curated sequence alignments to build powerful statistical models of sequence and structure conservation known as covariance models (CMs), implemented in the Infernal software package. We provide a step-by-step iterative protocol for identifying ncRNA homologs and then constructing an alignment and corresponding CM. We also work through an example for the bacterial small RNA MicA, discovering a previously unreported family of divergent MicA homologs in genus Xenorhabdus in the process. © 2016 by John Wiley & Sons, Inc. Keywords: covariance model; homology; RNA; Rfam; alignment; ncRNA; conservation

  • Rfam 12.0: updates to the RNA families database
    Nucleic Acids Research, 2014
    Co-Authors: Eric P. Nawrocki, Ruth Y Eberhardt, Alex Bateman, Sean R. Eddy, Paul P Gardner, Jennifer Daub, Sarah W. Burge, Evan W. Floden, Thomas A. Jones, John G Tate
    Abstract:

    The Rfam database (available at http://Rfam.xfam.org) is a collection of non-coding RNA families represented by manually curated sequence alignments, consensus secondary structures and annotation gathered from corresponding Wikipedia, taxonomy and ontology resources. In this article, we detail updates and improvements to the Rfam data and website for the Rfam 12.0 release. We describe the upgrade of our search pipeline to use Infernal 1.1 and demonstrate its improved homology detection ability by comparison with the previous version. The new pipeline is easier for users to apply to their own data sets, and we illustrate its ability to annotate RNAs in genomic and metagenomic data sets of various sizes. Rfam has been expanded to include 260 new families, including the well-studied large subunit ribosomal RNA family, and for the first time includes information on short sequence- and structure-based RNA motifs present within families.

  • Pacific Symposium on Biocomputing - Crowdsourcing RNA structural alignments with an online computer game.
    Biocomputing 2015, 2014
    Co-Authors: Jérôme Waldispühl, Arthur Kam, Paul P Gardner
    Abstract:

    The annotation and classification of ncRNAs is essential to decipher molecular mechanisms of gene regulation in normal and disease states. A database such as Rfam maintains alignments, consensus secondary structures, and corresponding annotations for RNA families. Its primary purpose is the automated, accurate annotation of non-coding RNAs in genomic sequences. However, the alignment of RNAs is computationally challenging, and the data stored in this database are often subject to improvements. Here, we design and evaluate Ribo, a human-computing game that aims to improve the accuracy of RNA alignments already stored in Rfam. We demonstrate the potential of our techniques and discuss the feasibility of large scale collaborative annotation and classification of RNA families.

  • Crowdsourcing RNA structural alignments with an online computer game
    2014
    Co-Authors: Jérôme Waldispühl, Arthur Kam, Paul P Gardner
    Abstract:

    The annotation and classification of ncRNAs is essential to decipher molecular mechanisms of gene regulation in normal and disease states. A database such as Rfam maintains alignments, consensus secondary structures, and corresponding annotations for RNA families. Its primary purpose is the automated, accurate annotation of non-coding RNAs in genomic sequences. However, the alignment of RNAs is computationally challenging, and the data stored in this database are often subject to improvements. Here, we design and evaluate Ribo, a human-computing game that aims to improve the accuracy of RNA alignments already stored in Rfam. We demonstrate the potential of our techniques and discuss the feasibility of large scale collaborative annotation and classification of RNA families.

  • An Introduction to RNA Databases
    Methods in Molecular Biology, 2013
    Co-Authors: Marc P. Hoeppner, Lars E. Barquist, Paul P Gardner
    Abstract:

    We present an introduction to RNA databases. The history and technology behind RNA databases are briefly discussed. We examine differing methods of data collection and curation and discuss their impact on both the scope and accuracy of the resulting databases. Finally, we demonstrate these principles through detailed examination of four leading RNA databases: Noncode, miRBase, Rfam, and SILVA.

Sean R. Eddy - One of the best experts on this subject based on the ideXlab platform.

  • the pfam protein families database in 2019
    Nucleic Acids Research, 2019
    Co-Authors: Sara Elgebali, Jaina Mistry, Simon C. Potter, Matloob Qureshi, Alex Bateman, Sean R. Eddy, Aurelien Luciani, Lorna Richardson, Gustavo A Salazar, Alfredo Smart
    Abstract:

    : The last few years have witnessed significant changes in Pfam (https://pfam.xfam.org). The number of families has grown substantially to a total of 17,929 in release 32.0. New additions have been coupled with efforts to improve existing families, including refinement of domain boundaries, their classification into Pfam clans, as well as their functional annotation. We recently began to collaborate with the RepeatsDB resource to improve the definition of tandem repeat families within Pfam. We carried out a significant comparison to the structural classification database, namely the Evolutionary Classification of Protein Domains (ECOD) that led to the creation of 825 new families based on their set of uncharacterized families (EUFs). Furthermore, we also connected Pfam entries to the Sequence Ontology (SO) through mapping of the Pfam type definitions to SO terms. Since Pfam has many community contributors, we recently enabled the linking between authorship of all Pfam entries with the corresponding authors' ORCID identifiers. This effectively permits authors to claim credit for their Pfam curation and link them to their ORCID record.

  • Rfam 13.0: shifting to a genome-centric resource for non-coding RNA families.
    Nucleic Acids Research, 2017
    Co-Authors: Ioanna Kalvari, Alex Bateman, Sean R. Eddy, Robert D. Finn, Eric P. Nawrocki, Joanna Argasinska, Natalia Quinones-olvera, Elena Rivas, Anton I. Petrov
    Abstract:

    The Rfam database is a collection of RNA families in which each family is represented by a multiple sequence alignment, a consensus secondary structure, and a covariance model. In this paper we introduce Rfam release 13.0, which switches to a new genome-centric approach that annotates a non-redundant set of reference genomes with RNA families. We describe new web interface features including faceted text search and R-scape secondary structure visualizations. We discuss a new literature curation workflow and a pipeline for building families based on RNAcentral. There are 236 new families in release 13.0, bringing the total number of families to 2687. The Rfam website is http://Rfam.org.

  • The Pfam protein families database: Towards a more sustainable future
    Nucleic Acids Research, 2016
    Co-Authors: Robert D. Finn, Penelope Coggill, Jaina Mistry, Simon C. Potter, Matloob Qureshi, Ruth Y Eberhardt, Alex L. Mitchell, Sean R. Eddy, Marco Punta, Amaia Sangrador-vegas
    Abstract:

    In the last two years the Pfam database (http://pfam.xfam.org) has undergone a substantial reorganisation to reduce the effort involved in making a release, thereby permitting more frequent releases. Arguably the most significant of these changes is that Pfam is now primarily based on the UniProtKB reference proteomes, with the counts of matched sequences and species reported on the website restricted to this smaller set. Building families on reference proteomes sequences brings greater stability, which decreases the amount of manual curation required to maintain them. It also reduces the number of sequences displayed on the website, whilst still providing access to many important model organisms. Matches to the full UniProtKB database are, however, still available and Pfam annotations for individual UniProtKB sequences can still be retrieved. Some Pfam entries (1.6%) which have no matches to reference proteomes remain; we are working with UniProt to see if sequences from them can be incorporated into reference proteomes. Pfam-B, the automatically-generated supplement to Pfam, has been removed. The current release (Pfam 29.0) includes 16 295 entries and 559 clans. The facility to view the relationship between families within a clan has been improved by the introduction of a new tool.

  • Rfam 12.0: updates to the RNA families database
    Nucleic Acids Research, 2014
    Co-Authors: Eric P. Nawrocki, Ruth Y Eberhardt, Alex Bateman, Sean R. Eddy, Paul P Gardner, Jennifer Daub, Sarah W. Burge, Evan W. Floden, Thomas A. Jones, John G Tate
    Abstract:

    The Rfam database (available at http://Rfam.xfam.org) is a collection of non-coding RNA families represented by manually curated sequence alignments, consensus secondary structures and annotation gathered from corresponding Wikipedia, taxonomy and ontology resources. In this article, we detail updates and improvements to the Rfam data and website for the Rfam 12.0 release. We describe the upgrade of our search pipeline to use Infernal 1.1 and demonstrate its improved homology detection ability by comparison with the previous version. The new pipeline is easier for users to apply to their own data sets, and we illustrate its ability to annotate RNAs in genomic and metagenomic data sets of various sizes. Rfam has been expanded to include 260 new families, including the well-studied large subunit ribosomal RNA family, and for the first time includes information on short sequence- and structure-based RNA motifs present within families.

  • Rfam 11.0: 10 years of RNA families
    Nucleic Acids Research, 2012
    Co-Authors: Sarah W. Burge, Ruth Y Eberhardt, Sean R. Eddy, Eric P. Nawrocki, Paul P Gardner, Jennifer Daub, John G Tate, Lars Barquist, Alex Bateman
    Abstract:

    The Rfam database (available via the website at http://Rfam.sanger.ac.uk and through our mirror at http://Rfam.janelia.org) is a collection of non-coding RNA families, primarily RNAs with a conserved RNA secondary structure, including both RNA genes and mRNA cis-regulatory elements. Each family is represented by a multiple sequence alignment, predicted secondary structure and covariance model. Here we discuss updates to the database in the latest release, Rfam 11.0, including the introduction of genome-based alignments for large families, the introduction of the Rfam Biomart as well as other user interface improvements. Rfam is available under the Creative Commons Zero license.

Alex Bateman - One of the best experts on this subject based on the ideXlab platform.

  • the pfam protein families database in 2019
    Nucleic Acids Research, 2019
    Co-Authors: Sara Elgebali, Jaina Mistry, Simon C. Potter, Matloob Qureshi, Alex Bateman, Sean R. Eddy, Aurelien Luciani, Lorna Richardson, Gustavo A Salazar, Alfredo Smart
    Abstract:

    : The last few years have witnessed significant changes in Pfam (https://pfam.xfam.org). The number of families has grown substantially to a total of 17,929 in release 32.0. New additions have been coupled with efforts to improve existing families, including refinement of domain boundaries, their classification into Pfam clans, as well as their functional annotation. We recently began to collaborate with the RepeatsDB resource to improve the definition of tandem repeat families within Pfam. We carried out a significant comparison to the structural classification database, namely the Evolutionary Classification of Protein Domains (ECOD) that led to the creation of 825 new families based on their set of uncharacterized families (EUFs). Furthermore, we also connected Pfam entries to the Sequence Ontology (SO) through mapping of the Pfam type definitions to SO terms. Since Pfam has many community contributors, we recently enabled the linking between authorship of all Pfam entries with the corresponding authors' ORCID identifiers. This effectively permits authors to claim credit for their Pfam curation and link them to their ORCID record.

  • Non‐Coding RNA Analysis Using the Rfam Database
    Current Protocols in Bioinformatics, 2018
    Co-Authors: Ioanna Kalvari, Alex Bateman, Robert D. Finn, Eric P. Nawrocki, Joanna Argasinska, Natalia Quinones-olvera, Anton I. Petrov
    Abstract:

    Rfam is a database of non-coding RNA families in which each family is represented by a multiple sequence alignment, a consensus secondary structure, and a covariance model. Using a combination of manual and literature-based curation and a custom software pipeline, Rfam converts descriptions of RNA families found in the scientific literature into computational models that can be used to annotate RNAs belonging to those families in any DNA or RNA sequence. Valuable research outputs that are often locked up in figures and supplementary information files are encapsulated in Rfam entries and made accessible through the Rfam Web site. The data produced by Rfam have a broad application, from genome annotation to providing training sets for algorithm development. This article gives an overview of how to search and navigate the Rfam Web site, and how to annotate sequences with RNA families. The Rfam database is freely available at http://Rfam.org. © 2018 by John Wiley & Sons, Inc.

  • non coding rna analysis using the Rfam database
    Current protocols in human genetics, 2018
    Co-Authors: Ioanna Kalvari, Alex Bateman, Robert D. Finn, Eric P. Nawrocki, Joanna Argasinska, Natalia Quinonesolvera, Anton I. Petrov
    Abstract:

    Rfam is a database of non-coding RNA families in which each family is represented by a multiple sequence alignment, a consensus secondary structure, and a covariance model. Using a combination of manual and literature-based curation and a custom software pipeline, Rfam converts descriptions of RNA families found in the scientific literature into computational models that can be used to annotate RNAs belonging to those families in any DNA or RNA sequence. Valuable research outputs that are often locked up in figures and supplementary information files are encapsulated in Rfam entries and made accessible through the Rfam Web site. The data produced by Rfam have a broad application, from genome annotation to providing training sets for algorithm development. This article gives an overview of how to search and navigate the Rfam Web site, and how to annotate sequences with RNA families. The Rfam database is freely available at http://Rfam.org. © 2018 by John Wiley & Sons, Inc.

  • Rfam 13.0: shifting to a genome-centric resource for non-coding RNA families.
    Nucleic Acids Research, 2017
    Co-Authors: Ioanna Kalvari, Alex Bateman, Sean R. Eddy, Robert D. Finn, Eric P. Nawrocki, Joanna Argasinska, Natalia Quinones-olvera, Elena Rivas, Anton I. Petrov
    Abstract:

    The Rfam database is a collection of RNA families in which each family is represented by a multiple sequence alignment, a consensus secondary structure, and a covariance model. In this paper we introduce Rfam release 13.0, which switches to a new genome-centric approach that annotates a non-redundant set of reference genomes with RNA families. We describe new web interface features including faceted text search and R-scape secondary structure visualizations. We discuss a new literature curation workflow and a pipeline for building families based on RNAcentral. There are 236 new families in release 13.0, bringing the total number of families to 2687. The Rfam website is http://Rfam.org.

  • Rfam 12.0: updates to the RNA families database
    Nucleic Acids Research, 2014
    Co-Authors: Eric P. Nawrocki, Ruth Y Eberhardt, Alex Bateman, Sean R. Eddy, Paul P Gardner, Jennifer Daub, Sarah W. Burge, Evan W. Floden, Thomas A. Jones, John G Tate
    Abstract:

    The Rfam database (available at http://Rfam.xfam.org) is a collection of non-coding RNA families represented by manually curated sequence alignments, consensus secondary structures and annotation gathered from corresponding Wikipedia, taxonomy and ontology resources. In this article, we detail updates and improvements to the Rfam data and website for the Rfam 12.0 release. We describe the upgrade of our search pipeline to use Infernal 1.1 and demonstrate its improved homology detection ability by comparison with the previous version. The new pipeline is easier for users to apply to their own data sets, and we illustrate its ability to annotate RNAs in genomic and metagenomic data sets of various sizes. Rfam has been expanded to include 260 new families, including the well-studied large subunit ribosomal RNA family, and for the first time includes information on short sequence- and structure-based RNA motifs present within families.