The Experts below are selected from a list of 12063 Experts worldwide ranked by ideXlab platform

David Roy Smith - One of the best experts on this subject based on the ideXlab platform.

  • The plastid genomes of nonphotosynthetic algae are not so small after all.
    Communicative & Integrative Biology, 2017
    Co-Authors: Francisco Figueroa-martinez, Aurora M. Nedelcu, Adrian Reyes-prieto, David Roy Smith
    Abstract:

    The thing about plastid genomes in nonphotosynthetic plants and algae is that they are usually very small and highly compact. This is not surprising: a heterotrophic existence means that genes for photosynthesis can be easily discarded. But the loss of photosynthesis cannot explain why the plastomes of heterotrophs are so often depauperate in Noncoding DNA. If plastid genomes from photosynthetic taxa can span the gamut of compactness, why can't those of nonphotosynthetic species? Well, recently we showed that they can. The free-living, heterotrophic green alga Polytoma uvella has a plastid genome boasting more than 165 kilobases of Noncoding DNA, making it the most bloated plastome yet found in a heterotroph. In this addendum to the primary study, we elaborate on why the P. uvella plastome is so inflated, discussing the potential impact of a free-living vs. parasitic lifestyle on plastid genome expansion in nonphotosynthetic lineages.

  • Amount Noncoding ptDNA regressed on plastid genome size with mode of inheritance indicated and amount Noncoding ptDNA regressed on plastid genome size with major taxonomic group indicated.
    2013
    Co-Authors: Kate Crosby, David Roy Smith
    Abstract:

    Dashed lines on both figures indicate the 25% and 75% bounds for percent of Noncoding DNA in a plastid genome. Analysis was carried out with all taxa (n = 82), and with logged variables. *We present the raw data here with Volvox carteri not pictured for ease of visual display (n = 81).

  • low nucleotide diversity for the expanded organelle and nuclear genomes of volvox carteri supports the mutational hazard hypothesis
    Molecular Biology and Evolution, 2010
    Co-Authors: David Roy Smith, Robert W Lee
    Abstract:

    The Noncoding-DNA content of organelle and nuclear genomes can vary immensely. Both adaptive and nonadaptive explanations for this variation have been proposed. This study addresses a nonadaptive explanation called the mutationalhazard hypothesis and applies it to the mitochondrial, plastid, and nuclear genomes of the multicellular green alga Volvox carteri. Given the expanded architecture of the V. carteri organelle and nuclear genomes (60–85% Noncoding DNA), the mutational-hazard hypothesis would predict them to have less silent-site nucleotide diversity (psilent) than their more compact counterparts from other eukaryotes—ultimately reflecting differences in 2Ngl (twice the effective number of genes per locus in the population times the mutation rate). The data presented here support this prediction: Analyses of mitochondrial, plastid, and nuclear DNAs from seven V. carteri forma nagariensis geographical isolates reveal low values of psilent (0.00038, 0.00065, and 0.00528, respectively), much lower values than those previously observed for the more compact organelle and nuclear DNAs of Chlamydomonas reinhardtii (a close relative of V. carteri). We conclude that the large Noncoding-DNA content of the V. carteri genomes is best explained by the mutational-hazard hypothesis and speculate that the shift from unicellular to multicellular life in the ancestor that gave rise to V. carteri contributed to a low V. carteri population size and thus a reduced 2Ngl. Complete mitochondrial and plastid genome maps for V. carteri are also presented and compared with those of C. reinhardtii.

  • The mitochondrial and plastid genomes of Volvox carteri: bloated molecules rich in repetitive DNA.
    BMC Genomics, 2009
    Co-Authors: David Roy Smith
    Abstract:

    Background The magnitude of Noncoding DNA in organelle genomes can vary significantly; it is argued that much of this variation is attributable to the dissemination of selfish DNA. The results of a previous study indicate that the mitochondrial DNA (mtDNA) of the green alga Volvox carteri abounds with palindromic repeats, which appear to be selfish elements. We became interested in the evolution and distribution of these repeats when, during a cursory exploration of the V. carteri nuclear DNA (nucDNA) and plastid DNA (ptDNA) sequences, we found palindromic repeats with similar structural features to those of the mtDNA. Upon this discovery, we decided to investigate the diversity and evolutionary implications of these palindromic elements by sequencing and characterizing large portions of mtDNA and ptDNA and then comparing these data to the V. carteri draft nuclear genome sequence.

Peter Andolfatto - One of the best experts on this subject based on the ideXlab platform.

  • a second generation assembly of the drosophila simulans genome provides new insights into patterns of lineage specific divergence
    Genome Research, 2013
    Co-Authors: Michael B Eisen, Kevin R Thornton, Peter Andolfatto
    Abstract:

    We create a new assembly of the Drosophila simulans genome using 142 million paired short-read sequences and previously published data for strain w(501). Our assembly represents a higher-quality genomic sequence with greater coverage, fewer misassemblies, and, by several indexes, fewer sequence errors. Evolutionary analysis of this genome reference sequence reveals interesting patterns of lineage-specific divergence that are different from those previously reported. Specifically, we find that Drosophila melanogaster evolves faster than D. simulans at all annotated classes of sites, including putatively neutrally evolving sites found in minimal introns. While this may be partly explained by a higher mutation rate in D. melanogaster, we also find significant heterogeneity in rates of evolution across classes of sites, consistent with historical differences in the effective population size for the two species. Also contrary to previous findings, we find that the X chromosome is evolving significantly faster than autosomes for nonsynonymous and most Noncoding DNA sites and significantly slower for synonymous sites. The absence of a X/A difference for putatively neutral sites and the robustness of the pattern to Gene Ontology and sex-biased expression suggest that partly recessive beneficial mutations may comprise a substantial fraction of Noncoding DNA divergence observed between species. Our results have more general implications for the interpretation of evolutionary analyses of genomes of different quality.

  • methods to detect selection on Noncoding DNA
    Methods of Molecular Biology, 2012
    Co-Authors: Ying Zhen, Peter Andolfatto
    Abstract:

    Vast tracts of Noncoding DNA contain elements that regulate gene expression in higher eukaryotes. Describing these regulatory elements and understanding how they evolve represent major challenges for biologists. Advances in the ability to survey genome-scale DNA sequence data are providing unprecedented opportunities to use evolutionary models and computational tools to identify functionally important elements and the mode of selection acting on them in multiple species. This chapter reviews some of the current methods that have been developed and applied on Noncoding DNA, what they have shown us, and how they are limited. Results of several recent studies reveal that a significantly larger fraction of Noncoding DNA in eukaryotic organisms is likely to be functional than previously believed, implying that the functional annotation of most Noncoding DNA in these organisms is largely incomplete. In Drosophila, recent studies have further suggested that a large fraction of Noncoding DNA divergence observed between species may be the product of recurrent adaptive substitution. Similar studies in humans have revealed a more complex pattern, with signatures of recurrent positive selection being largely concentrated in conserved Noncoding DNA elements. Understanding these patterns and the extent to which they generalize to other organisms awaits the analysis of forthcoming genome-scale polymorphism and divergence data from more species.

  • controlling type i error of the mcdonald kreitman test in genomewide scans for selection on Noncoding DNA
    Genetics, 2008
    Co-Authors: Peter Andolfatto
    Abstract:

    Departures from the assumption of homogenously interdigitated neutral and putatively selected sites in the McDonald–Kreitman test can lead to false rejections of the neutral model in the presence of intermediate levels of recombination. This problem is exacerbated by small sample sizes, nonequilibrium demography, recombination rate variation, and in comparisons involving more recently diverged species. I propose that establishing significance levels by coalescent simulation with recombination can improve the fidelity of the test in genomewide scans for selection on Noncoding DNA.

  • positive and negative selection on Noncoding DNA in drosophila simulans
    Molecular Biology and Evolution, 2008
    Co-Authors: Penelope R Haddrill, Doris Bachtrog, Peter Andolfatto
    Abstract:

    There is now a wealth of evidence that some of the most important regions of the genome are found outside those that encode proteins, and Noncoding regions of the genome have been shown to be subject to substantial levels of selective constraint, particularly in Drosophila. Recent work has suggested that these regions may also have been subject to the action of positive selection, with large fractions of Noncoding divergence having been driven to fixation by adaptive evolution. However, this work has focused on Drosophila melanogaster, which is thought to have experienced a reduction in effective population size (N(e)), and thus a reduction in the efficacy of selection, compared with its closest relative Drosophila simulans. Here, we examine patterns of evolution at several classes of Noncoding DNA in D. simulans and find that all Noncoding DNA is subject to the action of negative selection, indicated by reduced levels of polymorphism and divergence and a skew in the frequency spectrum toward rare variants. We find that the signature of negative selection on Noncoding DNA and nonsynonymous sites is obscured to some extent by purifying selection acting on preferred to unpreferred synonymous codon mutations. We investigate the extent to which divergence in Noncoding DNA is inferred to be the product of positive selection and to what extent these inferences depend on selection on synonymous sites and demography. Based on patterns of polymorphism and divergence for different classes of synonymous substitution, we find the divergence excess inferred in Noncoding DNA and nonsynonymous sites in the D. simulans lineage difficult to reconcile with demographic explanations.

  • selection recombination and demographic history in drosophila miranda
    Genetics, 2006
    Co-Authors: Doris Bachtrog, Peter Andolfatto
    Abstract:

    Selection, recombination, and the demographic history of a species can all have profound effects on genomewide patterns of variability. To assess the impact of these forces in the genome of Drosophila miranda, we examine polymorphism and divergence patterns at 62 loci scattered across the genome. In accordance with recent findings in D. melanogaster, we find that Noncoding DNA generally evolves more slowly than synonymous sites, that the distribution of polymorphism frequencies in Noncoding DNA is significantly skewed toward rare variants relative to synonymous sites, and that long introns evolve significantly slower than short introns or synonymous sites. These observations suggest that most Noncoding DNA is functionally constrained and evolving under purifying selection. However, in contrast to findings in the D. melanogaster species group, we find little evidence of adaptive evolution acting on either coding or Noncoding sequences in D. miranda. Levels of linkage disequilibrium (LD) in D. miranda are comparable to those observed in D. melanogaster, but vary considerably among chromosomes. These patterns suggest a significantly lower rate of recombination on autosomes, possibly due to the presence of polymorphic autosomal inversions and/or differences in chromosome sizes. All chromosomes show significant departures from the standard neutral model, including too much heterogeneity in synonymous site polymorphism relative to divergence among loci and a general excess of rare synonymous polymorphisms. These departures from neutral equilibrium expectations are discussed in the context of nonequilibrium models of demography and selection.

H. Eugene Stanley - One of the best experts on this subject based on the ideXlab platform.

  • species independence of mutual information in coding and Noncoding DNA
    Physical Review E, 2000
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, Hanspeter Herzel, H. Eugene Stanley
    Abstract:

    We explore if there exist universal statistical patterns that are different in coding and Noncoding DNA and can be found in all living organisms, regardless of their phylogenetic origin. We find that (i) the mutual information function [symbol: see text] has a significantly different functional form in coding and Noncoding DNA. We further find that (ii) the probability distributions of the average mutual information [symbol: see text] are significantly different in coding and Noncoding DNA, while (iii) they are almost the same for organisms of all taxonomic classes. Surprisingly, we find that [symbol: see text] is capable of predicting coding regions as accurately as organism-specific coding measures.

  • average mutual information of coding and Noncoding DNA
    Pacific Symposium on Biocomputing, 1999
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, H. Eugene Stanley, Dirk Holste, Hanspeter Herzel
    Abstract:

    One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and Noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and Noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and Noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

  • Long-range correlation properties of coding and Noncoding DNA sequences: GenBank analysis
    Physical Review E, 1995
    Co-Authors: Sergey V. Buldyrev, Martin Simons, C. K. Peng, Moshe E Matsa, Ary L Goldberger, Rosario N. Mantegna, Shlomo Havlin, H. Eugene Stanley
    Abstract:

    An open question in computational molecular biology is whether long-range correlations are present in both coding and Noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 Noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for Noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 Noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and Noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Sergey V. Buldyrev - One of the best experts on this subject based on the ideXlab platform.

  • species independence of mutual information in coding and Noncoding DNA
    Physical Review E, 2000
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, Hanspeter Herzel, H. Eugene Stanley
    Abstract:

    We explore if there exist universal statistical patterns that are different in coding and Noncoding DNA and can be found in all living organisms, regardless of their phylogenetic origin. We find that (i) the mutual information function [symbol: see text] has a significantly different functional form in coding and Noncoding DNA. We further find that (ii) the probability distributions of the average mutual information [symbol: see text] are significantly different in coding and Noncoding DNA, while (iii) they are almost the same for organisms of all taxonomic classes. Surprisingly, we find that [symbol: see text] is capable of predicting coding regions as accurately as organism-specific coding measures.

  • average mutual information of coding and Noncoding DNA
    Pacific Symposium on Biocomputing, 1999
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, H. Eugene Stanley, Dirk Holste, Hanspeter Herzel
    Abstract:

    One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and Noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and Noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and Noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

  • scaling features of Noncoding DNA
    Physica A-statistical Mechanics and Its Applications, 1999
    Co-Authors: H E Stanley, Ary L Goldberger, Sergey V. Buldyrev, Shlomo Havlin, Chungkang Peng, Michael Simons
    Abstract:

    We review evidence supporting the idea that the DNA sequence in genes containing Noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene, and utilize this fact to build a Coding Sequence Finder Algorithm, which uses statistical ideas to locate the coding regions of an unknown DNA sequence. Finally, we describe briefly some recent work adapting to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function, and reporting that Noncoding regions in eukaryotes display a larger redundancy than coding regions. Specifically, we consider the possibility that this result is solely a consequence of nucleotide concentration differences as first noted by Bonhoeffer and his collaborators. We find that cytosine-guanine (CG) concentration does have a strong "background" effect on redundancy. However, we find that for the purine-pyrimidine binary mapping rule, which is not affected by the difference in CG concentration, the Shannon redundancy for the set of analyzed sequences is larger for Noncoding regions compared to coding regions.

  • model of unequal chromosomal crossing over in DNA sequences
    Physica A-statistical Mechanics and Its Applications, 1998
    Co-Authors: Nikolay V Dokholyan, Sergey V. Buldyrev, Shlomo Havlin, Eugene H Stanley
    Abstract:

    It is known that some dimeric tandem repeats (DTR) are very abundant in Noncoding DNA. We find that certain DTR length distribution functions in Noncoding DNA can be fit by a power law function. We analyze a simplified model of unequal chromosomal crossing over and find that it produces a stable power law length distribution function, with the exponent μ=2. Although the exponent predicted by this model differs from those observed in nature, we argue that the biophysical process underlying this model provides the major contribution to the DTR length distribution function.

  • systematic analysis of coding and Noncoding DNA sequences using methods of statistical linguistics
    Physical Review E, 1995
    Co-Authors: Rosario N. Mantegna, Ary L Goldberger, Sergey V. Buldyrev, Shlomo Havlin, Chungkang Peng, Michael Simons, H E Stanley
    Abstract:

    We compare the statistical properties of coding and Noncoding regions in eukaryotic and viral DNA sequences by adapting two tests developed for the analysis of natural languages and symbolic sequences. The data set comprises all 30 sequences of length above 50 000 base pairs in GenBank Release No. 81.0, as well as the recently published sequences of C.elegans chromosome III (2.2 Mbp) and yeast chromosome XI (661 Kbp). We find that for the three chromosomes we studied the statistical properties of Noncoding regions appear to be closer to those observed in natural languages than those of the coding regions. In particular, (i) an n-tuple Zipf analysis of Noncoding regions reveals a regime close to power-law behavior while the coding regions show logarithmic behavior over a wide interval, while (ii) an n-gram entropy measurement shows that the Noncoding regions have a lower n-gram entropy (and hence a larger ``n-gram redundancy'') than the coding regions. In contrast to the three chromosomes, we find that for vertebrates\char22{}such as primates and rodents\char22{}and for viral DNA, the difference between the statistical properties of coding and Noncoding regions is not pronounced and therefore the results of the analyses of the investigated sequences are less conclusive. After noting the intrinsic limitations of the n-gram redundancy analysis, we also briefly discuss the failure of zero- and first-order Markovian models or simple nucleotide repeats to account fully for these ``linguistic'' features of DNA. Finally, we emphasize that our results by no means prove the existence of a ``language'' in Noncoding DNA.

Ivo Grosse - One of the best experts on this subject based on the ideXlab platform.

  • are Noncoding sequences of rickettsia prowazekii remnants of neutralized genes
    Journal of Molecular Evolution, 2000
    Co-Authors: Dirk Holste, Ivo Grosse, Olaf Weiss, Hanspeter Herzel
    Abstract:

    It has been hypothesized that a large fraction of 24% Noncoding DNA in R. prowazekii consists of degraded genes. This hypothesis has been based on the relatively high G+C content of Noncoding DNA. However, a comparison with other genomes also having a low overall G+C content shows that this argument would also apply to other bacteria. To test this hypothesis, we study the coding potential in sets of genes, pseudogenes, and intergenic regions. We find that the correlation function and the χ2-measure are clearly indicative of the coding function of genes and pseudogenes. However, both coding potentials make almost no indication of a preexisting reading frame in the remaining 23% of Noncoding DNA. We simulate the degradation of genes due to single-nucleotide substitutions and insertions/deletions and quantify the number of mutations required to remove indications of the reading frame. We discuss a reduced selection pressure as another possible origin of this comparatively large fraction of Noncoding sequences.

  • finding borders between coding and Noncoding DNA regions by an entropic segmentation method
    Physical Review Letters, 2000
    Co-Authors: Pedro Bernaolagalvan, Ivo Grosse, Pedro Carpena, Jose L Oliver, Ramon Romanroldan, Eugene H Stanley
    Abstract:

    We present a new computational approach to finding borders between coding and Noncoding DNA. This approach has two features: (i) DNA sequences are described by a 12-letter alphabet that captures the differential base composition at each codon position, and (ii) the search for the borders is carried out by means of an entropic segmentation method which uses only the general statistical properties of coding DNA. We find that this method is highly accurate in finding borders between coding and Noncoding regions and requires no iprior trainingi on known data sets. Our results appear to be more accurate than those obtained with moving windows in the discrimination of coding from Noncoding DNA. The entropic segmentation process partitions a heterogeneous DNA sequence into homogeneous subsequences, which we term compositional domains [1n3]. Although all the alphabets or mapping rules conventionally used in describing DNA sequences had been tried, we had not been able so far to assign any biological function to the obtained domains. Here we introduce a new alphabet that takes into account the differential base composition at each codon position, and we find that the compositional domains correlate to either coding or Noncoding DNA regions. This finding suggests the possibility of using the entropic segmentation process for computational gene finding. The computational recognition of genes is one of the challenges in the analysis of newly sequenced genomes, which is fundamental for modern functional genomics (the goal of which is the search for the different functional elements which make up the DNA sequences [4]).

  • species independence of mutual information in coding and Noncoding DNA
    Physical Review E, 2000
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, Hanspeter Herzel, H. Eugene Stanley
    Abstract:

    We explore if there exist universal statistical patterns that are different in coding and Noncoding DNA and can be found in all living organisms, regardless of their phylogenetic origin. We find that (i) the mutual information function [symbol: see text] has a significantly different functional form in coding and Noncoding DNA. We further find that (ii) the probability distributions of the average mutual information [symbol: see text] are significantly different in coding and Noncoding DNA, while (iii) they are almost the same for organisms of all taxonomic classes. Surprisingly, we find that [symbol: see text] is capable of predicting coding regions as accurately as organism-specific coding measures.

  • average mutual information of coding and Noncoding DNA
    Pacific Symposium on Biocomputing, 1999
    Co-Authors: Ivo Grosse, Sergey V. Buldyrev, H. Eugene Stanley, Dirk Holste, Hanspeter Herzel
    Abstract:

    One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and Noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and Noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and Noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.