The Experts below are selected from a list of 311796 Experts worldwide ranked by ideXlab platform

David Sankoff - One of the best experts on this subject based on the ideXlab platform.

  • PROCEEDINGS Open Access Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes
    2016
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Background: Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers’ embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. Results: We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Conclusions: Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomi

  • Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes.
    BMC bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomic syntenies provide appropriate input to a parameter-free procedure of multiple ortholog identification followed by Gene-Order reconstruction in solving instances of the "small phylogeny" problem.

  • Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes
    BMC Bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Abstract Background Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. Results We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Conclusions Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomic syntenies provide appropriate input to a parameter-free procedure of multiple ortholog identification followed by Gene-Order reconstruction in solving instances of the "small phylogeny" problem.

  • Gene Order in rosid phylogeny inferred from pairwise syntenies among extant genomes
    BMC Bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Background Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods.

  • ICCABS - Ancient angiosperm hexaploidy meets ancestral eudicot Gene Order
    2012 IEEE 2nd International Conference on Computational Advances in Bio and medical Sciences (ICCABS), 2012
    Co-Authors: Chunfang Zheng, Victor A. Albert, Eric Lyons, David Sankoff
    Abstract:

    We propose a protocol for reconstructing and analyzing the post-polyploidization ancestor of a set of genomes. Our method, applied to the post-hexaploid ancestor of six core eudicot flowering plants, reconstructs ancestral Gene Order, based on orthologs obtained for each pair of data genomes, harmonized into disjoint ortholog sets for multiple genomes.

Jijun Tang - One of the best experts on this subject based on the ideXlab platform.

  • Phylogeny analysis from Gene-Order data with massive duplications.
    BMC Genomics, 2017
    Co-Authors: Lingxi Zhou, Yu Lin, Bing Feng, Jieyi Zhao, Jijun Tang
    Abstract:

    Gene Order changes, under rearrangements, insertions, deletions and duplications, have been used as a new type of data source for phyloGenetic reconstruction. Because these changes are rare compared to sequence mutations, they allow the inference of phylogeny further back in evolutionary time. There exist many computational methods for the reconstruction of Gene-Order phylogenies, including widely used maximum parsimonious methods and maximum likelihood methods. However, both methods face challenges in handling large genomes with many duplicated Genes, especially in the presence of whole genome duplication. In this paper, we present three simple yet powerful methods based on maximum-likelihood (ML) approaches that encode multiplicities of both Gene adjacency and Gene content information for phyloGenetic reconstruction. Extensive experiments on simulated data sets show that our new method achieves the most accurate phylogenies compared to existing approaches. We also evaluate our method on real whole-genome data from eleven mammals. The package is publicly accessible at http://www.GeneOrder.org . Our new encoding schemes successfully incorporate the multiplicity information of Gene adjacencies and Gene content into an ML framework, and show promising results in reconstruct phylogenies for whole-genome data in the presence of massive duplications.

  • MLGO: phylogeny reconstruction and ancestral inference from Gene-Order data.
    BMC bioinformatics, 2014
    Co-Authors: Yu Lin, Jijun Tang
    Abstract:

    Background: The rapid accumulation of whole-genome data has renewed interest in the study of using Gene-Order data for phyloGenetic analyses and ancestral reconstruction. Current software and web servers typically do not support duplication and loss events along with rearrangements. Results: MLGO (Maximum Likelihood for Gene-Order Analysis) is a web tool for the reconstruction of phylogeny and/or ancestral genomes from Gene-Order data. MLGO is based on likelihood computation and shows advantages over existing methods in terms of accuracy, scalability and flexibility. Conclusions: To the best of our knowledge, it is the first web tool for analysis of large-scale genomic changes including not only rearrangements but also Gene insertions, deletions and duplications. The web tool is available from http://www.GeneOrder.org/server.php.

  • PhyloGenetic reconstruction with Gene Order and content information
    2011
    Co-Authors: Jijun Tang, William Arndt
    Abstract:

    PhyloGenetic reconstruction is the attempt to determine the evolutionary relationships which connect species by comparing their Genetic information. In recent decades research has opened a new frontier in this field, referred to as Gene Order and content data. This information format uses the Ordering and appearance of Genes on chromosomes to measure large scale evolutionary events and distances. The first problem addressed in this work is the inversion distance median problem, which involves finding a genome that minimizes the sum pairwise distance between itself and three other genomes. This median problem is known to be NP-hard and all existing solvers are extremely slow when genomes are distant. We present a new inversion median heuristic based on commuting reversals. Testing using simulated data sets shows that this method is a better trade-off between speed and accuracy than existing methods. Next addressed in this work is the current state of the art in Gene Order evolutionary measures is called the Double-Cut-and-Join distance. With the DCJ distance metric and a new concept called Prosthetic Chromosomes an elegant solution will be demonstrated for the situation where Gene Order data sets have unequal content. Egchel (Extended Gene Content HEueristic Layer) is our implementation which creates an equal content Gene Order data set to emulate the behavior of data sets which include insertion and deletion events. Testing of simulated data indicates that data sets which previously contained too few common Genes can now be analyzed using Egchel with practical speed and improved accuracy. Finally, Gene Order phyloGenetic analysis currently has the weakness of not having a convincing means to statistically validate trees. Existing literature has attempted to do so by resampling data sets with a jackknifing procedure adapted from sequence data analysis, but this approach has significant theoretical weakness. We have attempted to correct this weakness by incorporating a DCJ error model into a resampling method for verifying Gene Order based trees. These three methods work together to extend the speed, accuracy, and utility of phyloGenetic analysis using Gene Order and content based data.

  • CIBCB - Isolating - a new resampling method for Gene Order data
    2011 IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), 2011
    Co-Authors: Jian Shi, William Arndt, Jijun Tang
    Abstract:

    The purpose of using resampling methods on phyloGenetic data is to estimate the confidence value of branches. In recent years, bootstrapping and jackknifing are the two most popular resampling schemes which are widely used in biological reserach. However, for Gene Order data, traditional bootstrap procedures can not be applied because Gene Order data is viewed as one character with various states. Experience in the biological community has shown that jackknifing is a useful means of determining the confidence value of a Gene Order phylogeny. When genomes are distant, however, applying jackknifing tends to give low confidence values to many valid branches, causing them to be mistakenly removed. In this paper, we propose a new method that overcomes this disadvantage of jackknifing and achieves better accuracy and confidence values for Gene Order data. Compared to jackknifing, our experimental results show that the proposed method can produce phylogenies with lower error rates and much stronger support for good branches. We also establish a theoretic lower bound regarding how many Genes should be isolated, which is confirmed empirically.

  • CIBCB - Maximum likelihood phyloGenetic reconstruction using Gene Order encodings
    2011 IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), 2011
    Co-Authors: Nan Gao, Meng Zhang, Jijun Tang
    Abstract:

    Gene Order changes under rearrangement events such as inversions and transpositions have attracted increasing attention as a new type of data for phyloGenetic analysis. Since these events are rare, they allow the reconstruction of evolutionary history far back in time. Many software have been developed for the inference of Gene Order phylogenies, including widely used maximum parsimony methods such as GRAPPA and MGR. However, these methods confronted great difficulties in dealing with emerging large nuclear genomes. In this study, we proposed three simple yet powerful maximum likelihood(ML) based methods for phyloGenetic reconstruction by first encoding the Gene Orders into binary or multistate strings based on Gene adjacency information presented in the given genomes and further converting these strings into molecular sequences. RAxML is at last used to compute the maximum likelihood phylogeny. We conducted extensive experiments using simulated datasets and found that although the multistate encoding is more complex and more time-consuming, it did not improve accuracy over the methods using simpler binary encodings. Among all methods tested in our experiments, MLBE is of the most accuracy in most cases and often returns phylogenies without errors. ML methods is also fast and in the most difficult case only takes up to three days to compute datasets with 40 genomes, making it very suitable for large scale analysis. We give three simple and robust phyloGenetic reconstruction methods using different encodings based on maximum likelihood which has not been successfully applied for Gene Orderings before. Our development of these ML methods showed great potential in Gene Order analysis with respect to the high accuracy and stability, although formal mathematical and statistical analysis of these methods are much desired.

Laurence D. Hurst - One of the best experts on this subject based on the ideXlab platform.

  • Support for multiple classes of local expression clusters in Drosophila melanogaster, but no evidence for Gene Order conservation
    Genome biology, 2011
    Co-Authors: Claudia C. Weber, Laurence D. Hurst
    Abstract:

    Background Gene Order in eukaryotic genomes is not random, with Genes with similar expression profiles tending to cluster. In yeasts, the model taxon for Gene Order analysis, such syntenic clusters of non-homologous Genes tend to be conserved over evolutionary time. Whether similar clusters show Gene Order conservation in other lineages is, however, undecided. Here, we examine this issue in Drosophila melanogaster using high-resolution chromosome rearrangement data.

  • Support for multiple classes of local expression clusters in Drosophila melanogaster, but no evidence for Gene Order conservation
    Genome Biology, 2011
    Co-Authors: Claudia C. Weber, Laurence D. Hurst
    Abstract:

    Background Gene Order in eukaryotic genomes is not random, with Genes with similar expression profiles tending to cluster. In yeasts, the model taxon for Gene Order analysis, such syntenic clusters of non-homologous Genes tend to be conserved over evolutionary time. Whether similar clusters show Gene Order conservation in other lineages is, however, undecided. Here, we examine this issue in Drosophila melanogaster using high-resolution chromosome rearrangement data. Results We show that D. melanogaster has at least three classes of expression clusters: first, as observed in mammals, large clusters of functionally unrelated housekeeping Genes; second, small clusters of functionally related highly co-expressed Genes; and finally, as previously defined by Spellman and Rubin, larger domains of co-expressed but functionally unrelated Genes. The latter are, however, not independent of the small co-expression clusters and likely reflect a methodological artifact. While the small co-expression and housekeeping/essential Gene clusters resemble those observed in yeast, in contrast to yeast, we see no evidence that any of the three cluster types are preserved as synteny blocks. If anything, adjacent co-expressed Genes are more likely to become rearranged than expected. Again in contrast to yeast, in D. melanogaster , Gene pairs with short interGene distance or in divergent orientations tend to have higher rearrangement rates. These findings are consistent with co-expression being partly due to shared chromatin environment. Conclusions We conclude that, while similar in terms of cluster types, Gene Order evolution has strikingly different patterns in yeasts and in D. melanogaster , although recombination is associated with Gene Order rearrangement in both.

  • Stochasticity in Protein Levels Drives Colinearity of Gene Order in Metabolic Operons of Escherichia coli
    PLoS biology, 2009
    Co-Authors: Károly Kovács, Laurence D. Hurst, Balázs Papp
    Abstract:

    In bacterial genomes, Gene Order is not random. This is most evident when looking at operons, these often encoding enzymes involved in the same metabolic pathway or proteins from the same complex. Is Gene Order within operons nonrandom, however, and if so why? We examine this issue using metabolic operons as a case study. Using the metabolic network of Escherichia coli, we define the temporal Order of reactions. We find a pronounced trend for Genes to appear in operons in the same Order as they are needed in metabolism (colinearity). This is paradoxical as, at steady state, enzymes abundance should be independent of Order within the operon. We consider three extensions of the steady-state model that could potentially account for colinearity: (1) increased productivity associated with higher expression levels of the most 5′ Genes, (2) a faster metabolic processing immediately after up-regulation, and (3) metabolic stalling owing to stochastic protein loss. We establish the validity of these hypotheses by employing deterministic and stochastic models of enzyme kinetics. The stochastic stalling hypothesis correctly and uniquely predicts that colinearity is more pronounced both for lowly expressed operons and for Genes that are not physically adjacent. The alternative models fail to find any support. These results support the view that stochasticity is a pervasive problem to a cell and that Gene Order evolution can be driven by the selective consequences of fluctuations in protein levels.

  • Is optimal Gene Order impossible
    Trends in genetics : TIG, 2006
    Co-Authors: Juan F. Poyatos, Laurence D. Hurst
    Abstract:

    Recent evidence suggests that yeast Genes encoding proteins that are present in the same protein complex tend to be linked and to be co-expressed. More Generally, we found that Genes that are close to each other in the protein interaction network tend to be linked more often than expected and are often co-expressed. Unexpectedly, we found that linked Genes in network proximity have unusually high recombination rates. Because high recombination rates are associated with high rates of genome re-organization, our findings might explain why the clustering of Genes in proximity in the network is such a weak effect: there could be a co-evolutionary cycle of physical linkage for co-expression, upwards modification of the recombination rate and concomitant break-up of a cluster. Under such a model an ‘optimal' Gene Order is never stable.

  • The evolutionary dynamics of eukaryotic Gene Order
    Nature reviews. Genetics, 2004
    Co-Authors: Laurence D. Hurst, Csaba Pál, Martin J. Lercher
    Abstract:

    In eukaryotes, unlike in bacteria, Gene Order has typically been assumed to be random. However, the first statistically rigorous analyses of complete genomes, together with the availability of abundant Gene-expression data, have forced a paradigm shift: in every complete eukaryotic genome that has been analysed so far, Gene Order is not random. It seems that Genes that have similar and/or coordinated expression are often clustered. Here, we review this evidence and ask how such clusters evolve and how this relates to mechanisms that control Gene expression.

Lars Podsiadlowski - One of the best experts on this subject based on the ideXlab platform.

  • phylogeny and mitochondrial Gene Order variation in lophotrochozoa in the light of new mitogenomic data from nemertea
    BMC Genomics, 2009
    Co-Authors: Lars Podsiadlowski, Anke Braband, Jorn Von Dohren, Torsten H. Struck, Thomas Bartolomaeus
    Abstract:

    The new animal phylogeny established several taxa which were not identified by morphological analyses, most prominently the Ecdysozoa (arthropods, roundworms, priapulids and others) and Lophotrochozoa (molluscs, annelids, brachiopods and others). Lophotrochozoan interrelationships are under discussion, e.g. regarding the position of Nemertea (ribbon worms), which were discussed to be sister group to e.g. Mollusca, Brachiozoa or Platyhelminthes. Mitochondrial genomes contributed well with sequence data and Gene Order characters to the deep metazoan phylogeny debate. In this study we present the first complete mitochondrial genome record for a member of the Nemertea, Lineus viridis. Except two trnP and trnT, all Genes are located on the same strand. While Gene Order is most similar to that of the brachiopod Terebratulina retusa, sequence based analyses of mitochondrial Genes place nemerteans close to molluscs, phoronids and entoprocts without clear preference for one of these taxa as sister group. Almost all recent analyses with large datasets show good support for a taxon comprising Annelida, Mollusca, Brachiopoda, Phoronida and Nemertea. But the relationships among these taxa vary between different studies. The analysis of Gene Order differences gives evidence for a multiple independent occurrence of a large inversion in the mitochondrial genome of Lophotrochozoa and a re-inversion of the same part in gastropods. We hypothesize that some regions of the genome have a higher chance for intramolecular recombination than others and Gene Order data have to be analysed carefully to detect convergent rearrangement events.

  • Mitochondrial genome sequence and Gene Order of Sipunculus nudus give additional support for an inclusion of Sipuncula into Annelida
    BMC genomics, 2009
    Co-Authors: Adina Mwinyi, Thomas Bartolomaeus, Christoph Bleidorn, Achim Meyer, Bernhard Lieb, Lars Podsiadlowski
    Abstract:

    Mitochondrial genomes are a valuable source of data for analysing phyloGenetic relationships. Besides sequence information, mitochondrial Gene Order may add phyloGenetically useful information, too. Sipuncula are unsegmented marine worms, traditionally placed in their own phylum. Recent molecular and morphological findings suggest a close affinity to the segmented Annelida. The first complete mitochondrial genome of a member of Sipuncula, Sipunculus nudus, is presented. All 37 Genes characteristic for metazoan mtDNA were detected and are encoded on the same strand. The mitochondrial Gene Order (protein-coding and ribosomal RNA Genes) resembles that of annelids, but shows several derivations so far found only in Sipuncula. Sequence based phyloGenetic analysis of mitochondrial protein-coding Genes results in significant bootstrap support for Annelida sensu lato, combining Annelida together with Sipuncula, Echiura, Pogonophora and Myzostomida. The mitochondrial sequence data support a close relationship of Annelida and Sipuncula. Also the most parsimonious explanation of changes in Gene Order favours a derivation from the annelid Gene Order. These results complement findings from recent phyloGenetic analyses of nuclear encoded Genes as well as a report of a segmental neural patterning in Sipuncula.

  • Mitochondrial genome sequence and Gene Order of Sipunculus nudus give additional support for an inclusion of Sipuncula into Annelida
    BMC Genomics, 2009
    Co-Authors: Adina Mwinyi, Thomas Bartolomaeus, Christoph Bleidorn, Achim Meyer, Bernhard Lieb, Lars Podsiadlowski
    Abstract:

    Background Mitochondrial genomes are a valuable source of data for analysing phyloGenetic relationships. Besides sequence information, mitochondrial Gene Order may add phyloGenetically useful information, too. Sipuncula are unsegmented marine worms, traditionally placed in their own phylum. Recent molecular and morphological findings suggest a close affinity to the segmented Annelida. Results The first complete mitochondrial genome of a member of Sipuncula, Sipunculus nudus , is presented. All 37 Genes characteristic for metazoan mtDNA were detected and are encoded on the same strand. The mitochondrial Gene Order (protein-coding and ribosomal RNA Genes) resembles that of annelids, but shows several derivations so far found only in Sipuncula. Sequence based phyloGenetic analysis of mitochondrial protein-coding Genes results in significant bootstrap support for Annelida sensu lato , combining Annelida together with Sipuncula, Echiura, Pogonophora and Myzostomida. Conclusion The mitochondrial sequence data support a close relationship of Annelida and Sipuncula. Also the most parsimonious explanation of changes in Gene Order favours a derivation from the annelid Gene Order. These results complement findings from recent phyloGenetic analyses of nuclear encoded Genes as well as a report of a segmental neural patterning in Sipuncula.

Chunfang Zheng - One of the best experts on this subject based on the ideXlab platform.

  • PROCEEDINGS Open Access Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes
    2016
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Background: Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers’ embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. Results: We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Conclusions: Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomi

  • Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes.
    BMC bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomic syntenies provide appropriate input to a parameter-free procedure of multiple ortholog identification followed by Gene-Order reconstruction in solving instances of the "small phylogeny" problem.

  • Gene Order in rosid phylogeny, inferred from pairwise syntenies among extant genomes
    BMC Bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Abstract Background Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods. Results We address these problems, using the Gene Order of four core eudicot genomes (cacao, castor bean, papaya and grapevine) that have escaped any recent WGD events, and two others (poplar and cucumber) that descend from independent WGDs, in inferring the ancestral Gene Order of the rosid clade and those of its main subgroups, the fabids and malvids. We improve and adapt techniques including the OMG method for extracting large, paralogy-free, multiple orthologies from conflated pairwise synteny data among the six genomes and the PATHGROUPS approach for ancestral Gene Order reconstruction in a given phylogeny, where some genomes may be descendants of WGD events. We use the Gene Order evidence to evaluate the hypothesis that the Order Malpighiales belongs to the malvids rather than as traditionally assigned to the fabids. Conclusions Gene Orders of ancestral eudicot species, involving 10,000 or more Genes can be reconstructed in an efficient, parsimonious and consistent way, despite paralogies due to WGD and other processes. Pairwise genomic syntenies provide appropriate input to a parameter-free procedure of multiple ortholog identification followed by Gene-Order reconstruction in solving instances of the "small phylogeny" problem.

  • Gene Order in rosid phylogeny inferred from pairwise syntenies among extant genomes
    BMC Bioinformatics, 2012
    Co-Authors: Chunfang Zheng, David Sankoff
    Abstract:

    Background Ancestral Gene Order reconstruction for flowering plants has lagged behind developments in yeasts, insects and higher animals, because of the recency of widespread plant genome sequencing, sequencers' embargoes on public data use, paralogies due to whole genome duplication (WGD) and fractionation of undeleted duplicates, extensive paralogy from other sources, and the computational cost of existing methods.

  • ICCABS - Ancient angiosperm hexaploidy meets ancestral eudicot Gene Order
    2012 IEEE 2nd International Conference on Computational Advances in Bio and medical Sciences (ICCABS), 2012
    Co-Authors: Chunfang Zheng, Victor A. Albert, Eric Lyons, David Sankoff
    Abstract:

    We propose a protocol for reconstructing and analyzing the post-polyploidization ancestor of a set of genomes. Our method, applied to the post-hexaploid ancestor of six core eudicot flowering plants, reconstructs ancestral Gene Order, based on orthologs obtained for each pair of data genomes, harmonized into disjoint ortholog sets for multiple genomes.