The Experts below are selected from a list of 78792 Experts worldwide ranked by ideXlab platform

Eng-ti Leslie Low - One of the best experts on this subject based on the ideXlab platform.

  • Seqping: Gene Prediction pipeline for plant genomes using self-training Gene models and transcriptomic data
    BMC Bioinformatics, 2017
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Background Gene Prediction is one of the most important steps in the genome annotation process. A large number of software tools and pipelines developed by various computing techniques are available for Gene Prediction. However, these systems have yet to accurately predict all or even most of the protein-coding regions. Furthermore, none of the currently available Gene-finders has a universal Hidden Markov Model (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion.

  • Seqping: Gene Prediction pipeline for plant genomes using self-training Gene models and transcriptomic data.
    BMC bioinformatics, 2017
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Gene Prediction is one of the most important steps in the genome annotation process. A large number of software tools and pipelines developed by various computing techniques are available for Gene Prediction. However, these systems have yet to accurately predict all or even most of the protein-coding regions. Furthermore, none of the currently available Gene-finders has a universal Hidden Markov Model (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion. We present an automated Gene Prediction pipeline, Seqping that uses self-training HMM models and transcriptomic data. The pipeline processes the genome and transcriptome sequences of the target species using GlimmerHMM, SNAP, and AUGUSTUS pipelines, followed by MAKER2 program to combine Predictions from the three tools in association with the transcriptomic evidence. Seqping Generates species-specific HMMs that are able to offer unbiased Gene Predictions. The pipeline was evaluated using the Oryza sativa and Arabidopsis thaliana genomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis showed that the pipeline was able to identify at least 95% of BUSCO's plantae dataset. Our evaluation shows that Seqping was able to Generate better Gene Predictions compared to three HMM-based programs (MAKER2, GlimmerHMM and AUGUSTUS) using their respective available HMMs. Seqping had the highest accuracy in rice (0.5648 for CDS, 0.4468 for exon, and 0.6695 nucleotide structure) and A. thaliana (0.5808 for CDS, 0.5955 for exon, and 0.8839 nucleotide structure). Seqping provides researchers a seamless pipeline to train species-specific HMMs and predict Genes in newly sequenced or less-studied genomes. We conclude that the Seqping pipeline Predictions are more accurate than Gene Predictions using the other three approaches with the default or available HMMs.

  • Seqping: Gene Prediction Pipeline for Plant Genomes using Self- Trained Gene Models and Transcriptomic Data
    2016
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Although various software are available for Gene Prediction, none of the currently available Gene-finders have a universal Hidden Markov Models (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion. Here, we report an automated pipeline that performs Gene Prediction using self-trained HMM models and transcriptomic data. The program processes the genome and transcriptome sequences of a target species through GlimmerHMM, SNAP, and AUGUSTUS training pipeline that ends with the program MAKER2 combining the Predictions from the three models in association with the transcriptomic evidence. The pipeline Generates species-specific HMMs and is able to predict Genes that are not biased to other model organisms. Our evaluation of the program revealed that it performed better than the use of the closest related HMM from a standalone program.

Michael R Brent - One of the best experts on this subject based on the ideXlab platform.

  • Using ESTs to improve the accuracy of de novo Gene Prediction
    BMC bioinformatics, 2006
    Co-Authors: Chaochun Wei, Michael R Brent
    Abstract:

    ESTs are a tremendous resource for determining the exon-intron structures of Genes, but even extensive EST sequencing tends to leave many exons and Genes untouched. Gene Prediction systems based exclusively on EST alignments miss these exons and Genes, leading to poor sensitivity. De novo Gene Prediction systems, which ignore ESTs in favor of genomic sequence, can predict such "untouched" exons, but they are less accurate when predicting exons to which ESTs align. TWINSCAN is the most accurate de novo Gene finder available for nematodes and N-SCAN is the most accurate for mammals, as measured by exact CDS Gene Prediction and exact exon Prediction. TWINSCAN_EST is a new system that successfully combines EST alignments with TWINSCAN. On the whole C. elegans genome TWINSCAN_EST shows 14% improvement in sensitivity and 13% in specificity in predicting exact Gene structures compared to TWINSCAN without EST alignments. Not only are the structures revealed by EST alignments predicted correctly, but these also constrain the Predictions without alignments, improving their accuracy. For the human genome, we used the same approach with N-SCAN, creating N-SCAN_EST. On the whole genome, N-SCAN_EST produced a 6% improvement in sensitivity and 1% in specificity of exact Gene structure Predictions compared to N-SCAN. TWINSCAN_EST and N-SCAN_EST are more accurate than TWINSCAN and N-SCAN, while retaining their ability to discover novel Genes to which no ESTs align. Thus, we recommend using the EST versions of these programs to annotate any genome for which EST information is available. TWINSCAN_EST and N-SCAN_EST are part of the TWINSCAN open source software package http://Genes.cse.wustl.edu/distribution/download_TS.html .

  • iterative Gene Prediction and pseudoGene removal improves genome annotation
    Genome Research, 2006
    Co-Authors: Marijke J Van Baren, Michael R Brent
    Abstract:

    Correct Gene Prediction is impaired by the presence of processed pseudoGenes: nonfunctional, intronless copies of real Genes found elsewhere in the genome. Gene Prediction programs frequently mistake processed pseudoGenes for real Genes or exons, leading to biologically irrelevant Gene Predictions. While methods exist to identify processed pseudoGenes in genomes, no attempt has been made to integrate pseudoGene removal with Gene Prediction, or even to provide a freestanding tool that identifies such erroneous Gene Predictions. We have created PPFINDER (for Processed PseudoGene finder), a program that integrates several methods of processed pseudoGene finding in mammalian Gene annotations. We used PPFINDER to remove pseudoGenes from N-SCAN Gene Predictions, and show that Gene Prediction improves substantially when Gene Prediction and pseudoGene masking are interleaved. In addition, we used PPFINDER with Gene Predictions as a parent database, eliminating the need for libraries of known Genes. This allows us to run the Gene Prediction/PPFINDER procedure on newly sequenced genomes for which few Genes are known.

  • iterative Gene Prediction and pseudoGene removal improves genome annotation
    Genome Research, 2006
    Co-Authors: Marijke J Van Baren, Michael R Brent
    Abstract:

    Correct Gene Prediction is impaired by the presence of processed pseudoGenes: nonfunctional, intronless copies of real Genes found elsewhere in the genome. Gene Prediction programs frequently mistake processed pseudoGenes for real Genes or exons, leading to biologically irrelevant Gene Predictions. While methods exist to identify processed pseudoGenes in genomes, no attempt has been made to integrate pseudoGene removal with Gene Prediction, or even to provide a freestanding tool that identifies such erroneous Gene Predictions. We have created PPFINDER (for Processed PseudoGene finder), a program that integrates several methods of processed pseudoGene finding in mammalian Gene annotations. We used PPFINDER to remove pseudoGenes from N-SCAN Gene Predictions, and show that Gene Prediction improves substantially when Gene Prediction and pseudoGene masking are interleaved. In addition, we used PPFINDER with Gene Predictions as a parent database, eliminating the need for libraries of known Genes. This allows us to run the Gene Prediction/PPFINDER procedure on newly sequenced genomes for which few Genes are known.

  • using multiple alignments to improve Gene Prediction
    Journal of Computational Biology, 2006
    Co-Authors: Samuel S Gross, Michael R Brent
    Abstract:

    The multiple species de novo Gene Prediction problem can be stated as follows: given an alignment of genomic sequences from two or more organisms, predict the location and structure of all protein-coding Genes in one or more of the sequences. Here, we present a new system, N-SCAN (a.k.a. TWINSCAN 3.0), for addressing this problem. N-SCAN can model the phyloGenetic relationships between the aligned genome sequences, context dependent substitution rates, and insertions and deletions. An implementation of N-SCAN was created and used to Generate Predictions for the entire human genome and the genome of the fruit fly Drosophila melanogaster. Analyses of the Predictions reveal that N-SCAN's accuracy in both human and fly exceeds that of all previously published whole-genome de novo Gene predictors.

  • using multiple alignments to improve Gene Prediction
    Research in Computational Molecular Biology, 2005
    Co-Authors: Samuel S Gross, Michael R Brent
    Abstract:

    The multiple species de novo Gene Prediction problem can be stated as follows: given an alignment of genomic sequences from two or more organisms, predict the location and structure of all protein-coding Genes in one or more of the sequences. Here, we present a new system, N-SCAN (a.k.a. TWINSCAN 3.0), for addressing this problem. N-SCAN has the ability to model dependencies between the aligned sequences, context-dependent substitution rates, and insertions and deletions in the sequences. An implementation of N-SCAN was created and used to Generate Predictions for the entire human genome. An analysis of the Predictions reveals that N-SCAN's predictive accuracy in human exceeds that of all previously published whole-genome de novo Gene predictors. In addition, Predictions were Generated for the genome of the fruit fly Drosophila melanogaster to demonstrate the applicability of N-SCAN to invertebrate Gene Prediction.

Kuang-lim Chan - One of the best experts on this subject based on the ideXlab platform.

  • Seqping: Gene Prediction pipeline for plant genomes using self-training Gene models and transcriptomic data
    BMC Bioinformatics, 2017
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Background Gene Prediction is one of the most important steps in the genome annotation process. A large number of software tools and pipelines developed by various computing techniques are available for Gene Prediction. However, these systems have yet to accurately predict all or even most of the protein-coding regions. Furthermore, none of the currently available Gene-finders has a universal Hidden Markov Model (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion.

  • Seqping: Gene Prediction pipeline for plant genomes using self-training Gene models and transcriptomic data.
    BMC bioinformatics, 2017
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Gene Prediction is one of the most important steps in the genome annotation process. A large number of software tools and pipelines developed by various computing techniques are available for Gene Prediction. However, these systems have yet to accurately predict all or even most of the protein-coding regions. Furthermore, none of the currently available Gene-finders has a universal Hidden Markov Model (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion. We present an automated Gene Prediction pipeline, Seqping that uses self-training HMM models and transcriptomic data. The pipeline processes the genome and transcriptome sequences of the target species using GlimmerHMM, SNAP, and AUGUSTUS pipelines, followed by MAKER2 program to combine Predictions from the three tools in association with the transcriptomic evidence. Seqping Generates species-specific HMMs that are able to offer unbiased Gene Predictions. The pipeline was evaluated using the Oryza sativa and Arabidopsis thaliana genomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis showed that the pipeline was able to identify at least 95% of BUSCO's plantae dataset. Our evaluation shows that Seqping was able to Generate better Gene Predictions compared to three HMM-based programs (MAKER2, GlimmerHMM and AUGUSTUS) using their respective available HMMs. Seqping had the highest accuracy in rice (0.5648 for CDS, 0.4468 for exon, and 0.6695 nucleotide structure) and A. thaliana (0.5808 for CDS, 0.5955 for exon, and 0.8839 nucleotide structure). Seqping provides researchers a seamless pipeline to train species-specific HMMs and predict Genes in newly sequenced or less-studied genomes. We conclude that the Seqping pipeline Predictions are more accurate than Gene Predictions using the other three approaches with the default or available HMMs.

  • Seqping: Gene Prediction Pipeline for Plant Genomes using Self- Trained Gene Models and Transcriptomic Data
    2016
    Co-Authors: Kuang-lim Chan, Rozana Rosli, Tatiana V. Tatarinova, Michael Hogan, Mohd Firdaus-raih, Eng-ti Leslie Low
    Abstract:

    Although various software are available for Gene Prediction, none of the currently available Gene-finders have a universal Hidden Markov Models (HMM) that can perform Gene Prediction for all organisms equally well in an automatic fashion. Here, we report an automated pipeline that performs Gene Prediction using self-trained HMM models and transcriptomic data. The program processes the genome and transcriptome sequences of a target species through GlimmerHMM, SNAP, and AUGUSTUS training pipeline that ends with the program MAKER2 combining the Predictions from the three models in association with the transcriptomic evidence. The pipeline Generates species-specific HMMs and is able to predict Genes that are not biased to other model organisms. Our evaluation of the program revealed that it performed better than the use of the closest related HMM from a standalone program.

Jan Grau - One of the best experts on this subject based on the ideXlab platform.

  • Combining RNA-seq data and homology-based Gene Prediction for plants, animals and fungi.
    BMC bioinformatics, 2018
    Co-Authors: Jens Keilwagen, Sven Twardziok, Frank Hartung, Michael Paulini, Jan Grau
    Abstract:

    Genome annotation is of key importance in many research questions. The identification of protein-coding Genes is often based on transcriptome sequencing data, ab-initio or homology-based Prediction. Recently, it was demonstrated that intron position conservation improves homology-based Gene Prediction, and that experimental data improves ab-initio Gene Prediction. Here, we present an extension of the Gene Prediction program GeMoMa that utilizes amino acid sequence conservation, intron position conservation and optionally RNA-seq data for homology-based Gene Prediction. We show on published benchmark data for plants, animals and fungi that GeMoMa performs better than the Gene Prediction programs BRAKER1, MAKER2, and CodingQuarry, and purely RNA-seq-based pipelines for transcript identification. In addition, we demonstrate that using multiple reference organisms may help to further improve the performance of GeMoMa. Finally, we apply GeMoMa to four nematode species and to the recently published barley reference genome indicating that current annotations of protein-coding Genes may be refined using GeMoMa Predictions. GeMoMa might be of great utility for annotating newly sequenced genomes but also for finding homologs of a specific Gene or Gene family. GeMoMa has been published under GNU GPL3 and is freely available at http://www.jstacs.de/index.php/GeMoMa .

  • Combining RNA-seq data and homology-based Gene Prediction for plants, animals and fungi
    2017
    Co-Authors: Jens Keilwagen, Sven Twardziok, Frank Hartung, Michael Paulini, Jan Grau
    Abstract:

    Motivation: Genome annotation is of key importance in many research questions. The identification of protein-coding Genes is often based on transcriptome sequencing data, ab-initio or homology-based Prediction. Recently, it was demonstrated that intron position conservation improves homology-based Gene Prediction, and that experimental data improves ab-initio Gene Prediction. Results: Here, we present an extension of the Gene Prediction tool GeMoMa that utilizes amino acid sequence conservation, intron position conservation and optionally RNA-seq data for homology-based Gene Prediction. We show on published benchmark data for plants, animals and fungi that GeMoMa performs better than the Gene Prediction programs BRAKER1, MAKER2, and CodingQuarry, and purely RNA-seq-based pipelines for transcript identification. In addition, we demonstrate that using multiple reference organisms may help to further improve the performance of GeMoMa. Finally, we apply GeMoMa to four nematode species and to the recently published barley reference genome indicating that current annotations of protein-coding Genes may be refined using GeMoMa Predictions. Availability: GeMoMa has been published under GNU GPL3 and is freely available at http://www.jstacs.de/index.php/GeMoMa.

  • Using intron position conservation for homology-based Gene Prediction
    Nucleic acids research, 2016
    Co-Authors: Jens Keilwagen, Michael Wenk, Jessica L. Erickson, Martin H. Schattat, Jan Grau, Frank Hartung
    Abstract:

    Annotation of protein-coding Genes is very important in bioinformatics and biology and has a decisive influence on many downstream analyses. Homology-based Gene Prediction programs allow for transferring knowledge about protein-coding Genes from an annotated organism to an organism of interest.Here, we present a homology-based Gene Prediction program called GeMoMa. GeMoMa utilizes the conservation of intron positions within Genes to predict related Genes in other organisms. We assess the performance of GeMoMa and compare it with state-of-the-art competitors on plant and animal genomes using an extended best reciprocal hit approach. We find that GeMoMa often makes more precise Predictions than its competitors yielding a substantially increased number of correct transcripts. Subsequently, we exemplarily validate GeMoMa Predictions using Sanger sequencing. Finally, we use RNA-seq data to compare the Predictions of homology-based Gene Prediction programs, and find again that GeMoMa performs well.Hence, we conclude that exploiting intron position conservation improves homology-based Gene Prediction, and we make GeMoMa freely available as command-line tool and Galaxy integration.

Jens Keilwagen - One of the best experts on this subject based on the ideXlab platform.

  • Combining RNA-seq data and homology-based Gene Prediction for plants, animals and fungi.
    BMC bioinformatics, 2018
    Co-Authors: Jens Keilwagen, Sven Twardziok, Frank Hartung, Michael Paulini, Jan Grau
    Abstract:

    Genome annotation is of key importance in many research questions. The identification of protein-coding Genes is often based on transcriptome sequencing data, ab-initio or homology-based Prediction. Recently, it was demonstrated that intron position conservation improves homology-based Gene Prediction, and that experimental data improves ab-initio Gene Prediction. Here, we present an extension of the Gene Prediction program GeMoMa that utilizes amino acid sequence conservation, intron position conservation and optionally RNA-seq data for homology-based Gene Prediction. We show on published benchmark data for plants, animals and fungi that GeMoMa performs better than the Gene Prediction programs BRAKER1, MAKER2, and CodingQuarry, and purely RNA-seq-based pipelines for transcript identification. In addition, we demonstrate that using multiple reference organisms may help to further improve the performance of GeMoMa. Finally, we apply GeMoMa to four nematode species and to the recently published barley reference genome indicating that current annotations of protein-coding Genes may be refined using GeMoMa Predictions. GeMoMa might be of great utility for annotating newly sequenced genomes but also for finding homologs of a specific Gene or Gene family. GeMoMa has been published under GNU GPL3 and is freely available at http://www.jstacs.de/index.php/GeMoMa .

  • Combining RNA-seq data and homology-based Gene Prediction for plants, animals and fungi
    2017
    Co-Authors: Jens Keilwagen, Sven Twardziok, Frank Hartung, Michael Paulini, Jan Grau
    Abstract:

    Motivation: Genome annotation is of key importance in many research questions. The identification of protein-coding Genes is often based on transcriptome sequencing data, ab-initio or homology-based Prediction. Recently, it was demonstrated that intron position conservation improves homology-based Gene Prediction, and that experimental data improves ab-initio Gene Prediction. Results: Here, we present an extension of the Gene Prediction tool GeMoMa that utilizes amino acid sequence conservation, intron position conservation and optionally RNA-seq data for homology-based Gene Prediction. We show on published benchmark data for plants, animals and fungi that GeMoMa performs better than the Gene Prediction programs BRAKER1, MAKER2, and CodingQuarry, and purely RNA-seq-based pipelines for transcript identification. In addition, we demonstrate that using multiple reference organisms may help to further improve the performance of GeMoMa. Finally, we apply GeMoMa to four nematode species and to the recently published barley reference genome indicating that current annotations of protein-coding Genes may be refined using GeMoMa Predictions. Availability: GeMoMa has been published under GNU GPL3 and is freely available at http://www.jstacs.de/index.php/GeMoMa.

  • Using intron position conservation for homology-based Gene Prediction
    Nucleic acids research, 2016
    Co-Authors: Jens Keilwagen, Michael Wenk, Jessica L. Erickson, Martin H. Schattat, Jan Grau, Frank Hartung
    Abstract:

    Annotation of protein-coding Genes is very important in bioinformatics and biology and has a decisive influence on many downstream analyses. Homology-based Gene Prediction programs allow for transferring knowledge about protein-coding Genes from an annotated organism to an organism of interest.Here, we present a homology-based Gene Prediction program called GeMoMa. GeMoMa utilizes the conservation of intron positions within Genes to predict related Genes in other organisms. We assess the performance of GeMoMa and compare it with state-of-the-art competitors on plant and animal genomes using an extended best reciprocal hit approach. We find that GeMoMa often makes more precise Predictions than its competitors yielding a substantially increased number of correct transcripts. Subsequently, we exemplarily validate GeMoMa Predictions using Sanger sequencing. Finally, we use RNA-seq data to compare the Predictions of homology-based Gene Prediction programs, and find again that GeMoMa performs well.Hence, we conclude that exploiting intron position conservation improves homology-based Gene Prediction, and we make GeMoMa freely available as command-line tool and Galaxy integration.