The Experts below are selected from a list of 131418 Experts worldwide ranked by ideXlab platform

Ying Zhou - One of the best experts on this subject based on the ideXlab platform.

  • length distribution of ancestral tracks under a general admixture model and its applications in population history inference
    Scientific Reports, 2016
    Co-Authors: Xiong Yang, Wei Guo, Kai Yuan, Ying Zhou
    Abstract:

    The length of ancestral tracks decays with the passing of generations which can be used to infer population admixture histories. Previous studies have shown the power in recovering the histories of admixed populations via the length distributions of ancestral tracks even under simple models. We believe that the deduction of length distributions under a general model will greatly elevate the power. Here we first deduced the length distributions under a general model and proposed general principles in parameter estimation and model selection with the deduced length distributions. Next, we focused on studying the length distributions and its applications under three typical special cases. Extensive simulations showed that the length distributions of ancestral tracks were well predicted by our theoretical framework. We further developed a new method, AdmixInfer, based on the length distributions and good performance was observed when it was applied to infer population histories under the three typical models. Notably, our method was insensitive to demographic history, sample size and threshold to discard short tracks. Finally, good performance was also observed when applied to some real datasets of African Americans, Mexicans and South Asian populations from the HapMap Project and the Human Genome Diversity Project.

  • length distribution of ancestral tracks under a general admixture model and its applications in population history inference
    bioRxiv, 2015
    Co-Authors: Xiong Yang, Wei Guo, Kai Yuan, Ying Zhou
    Abstract:

    As a chromosome is sliced into pieces by recombination after entering an admixed population, ancestral tracks of chromosomes are shortened with the pasting of generations. The length distribution of ancestral tracks reflects information of recombination and thus can be used to infer the histories of admixed populations. Previous studies have shown that inference based on ancestral tracks is powerful in recovering the histories of admixed populations. However, population histories are always complex, and previous studies only deduced the length distribution of ancestral tracks under very simple admixture models. The deduction of length distribution of ancestral tracks under a more general model will greatly elevate the power in inferring population histories. Here we first deduced the length distribution of ancestral tracks under a general model in an admixed population, and proposed general principles in parameter estimation and model selection with the length distribution. Next, we focused on studying the length distribution of ancestral tracks and its applications under three typical admixture models, which were all special cases of our general model. Extensive simulations showed that the length distribution of ancestral tracks was well predicted by our theoretical models. We further developed a new method based on the length distribution of ancestral tracks and good performance was observed when it was applied in inferring population histories under the three typical models. Notably, our method was insensitive to demographic history, sample size and threshold to discard short tracks. Finally, we applied our method in African Americans and Mexicans from the HapMap dataset, and several South Asian populations from the Human Genome Diversity Project dataset. The results showed that the histories of African Americans and Mexicans matched the historical records well, and the population admixture history of South Asians was very complex and could be traced back to around 100 generations ago.

L L Cavallisforza - One of the best experts on this subject based on the ideXlab platform.

  • the human genome Diversity Project past present and future
    Nature Reviews Genetics, 2005
    Co-Authors: L L Cavallisforza
    Abstract:

    The Human Genome Project, in accomplishing its goal of sequencing one human genome, heralded a new era of research, a component of which is the systematic study of human genetic variation. Despite delays, the Human Genome Diversity Project has started to make progress in understanding the patterns of this variation and its causes, and also promises to provide important information for biomedical studies.

  • the chinese human genome Diversity Project
    Proceedings of the National Academy of Sciences of the United States of America, 1998
    Co-Authors: L L Cavallisforza
    Abstract:

    The Chinese population comprises one-fifth of the human species. The Chinese government officially recognizes 56 ethnic groups, one of which is the Han majority (1 billion and 100 million people), and the other 55 are ethnic minorities (totaling about 100 million). The latter are spread over most of China, but especially in the south. Close to half of the minorities are found in one of the 28 provinces of China, Yunnan. The distinction is primarily linguistic but corresponds closely to other cultural differences. The paper by Chu et al . published in this issue of the Proceedings (1) explores the genetic stratification of about half of the official ethnic subdivisions by means of microsatellites, a class of genetic markers recently discovered that has proved very useful for several purposes. The paper represents the collective effort of several institutes participating in the Chinese Human Genome Diversity Project (CHGDP). The broader Human Genome Diversity Project (HGDP) was generated in 1991 by the international Human Genome Organization (HUGO) and is regionally organized (see http://www.stanford.edu/group/morrinst/HGDP/html). The CHGDP has started collecting cell lines from the official ethnic groups and testing their DNAs. The 56 official ethnic groups do not exhaust current Chinese Diversity, as there are more than 100 languages spoken in China, but they include the most important ones. Microsatellites are repeats of short DNA segments, practically less than five nucleotides long. They have a high mutation rate and therefore a large number of alleles, which makes them perhaps three times more informative on average than the most common type of genetic polymorphisms, single nucleotide substitutions, which are mostly biallelic. They are used very widely in genetic linkage studies and have begun to be used in evolutionary analyses (e.g., refs. 2–4). Thirty microsatellites were tested by Chu et al. (1) for reconstructing …

Swapan Mallick - One of the best experts on this subject based on the ideXlab platform.

  • the simons genome Diversity Project 300 genomes from 142 diverse populations
    Nature, 2016
    Co-Authors: Swapan Mallick, Heng Li, Mark Lipson, Iain Mathieson, Melissa Gymrek, Fernando Racimo, Mengyao Zhao
    Abstract:

    Here we report the Simons Genome Diversity Project data set: high quality genomes from 300 individuals from 142 diverse populations. These genomes include at least 5.8 million base pairs that are not present in the human reference genome. Our analysis reveals key features of the landscape of human genome variation, including that the rate of accumulation of mutations has accelerated by about 5% in non-Africans compared to Africans since divergence. We show that the ancestors of some pairs of present-day human populations were substantially separated by 100,000 years ago, well before the archaeologically attested onset of behavioural modernity. We also demonstrate that indigenous Australians, New Guineans and Andamanese do not derive substantial ancestry from an early dispersal of modern humans; instead, their modern human ancestry is consistent with coming from the same source as that of other non-Africans.

Mengyao Zhao - One of the best experts on this subject based on the ideXlab platform.

  • the simons genome Diversity Project 300 genomes from 142 diverse populations
    Nature, 2016
    Co-Authors: Swapan Mallick, Heng Li, Mark Lipson, Iain Mathieson, Melissa Gymrek, Fernando Racimo, Mengyao Zhao
    Abstract:

    Here we report the Simons Genome Diversity Project data set: high quality genomes from 300 individuals from 142 diverse populations. These genomes include at least 5.8 million base pairs that are not present in the human reference genome. Our analysis reveals key features of the landscape of human genome variation, including that the rate of accumulation of mutations has accelerated by about 5% in non-Africans compared to Africans since divergence. We show that the ancestors of some pairs of present-day human populations were substantially separated by 100,000 years ago, well before the archaeologically attested onset of behavioural modernity. We also demonstrate that indigenous Australians, New Guineans and Andamanese do not derive substantial ancestry from an early dispersal of modern humans; instead, their modern human ancestry is consistent with coming from the same source as that of other non-Africans.

Chris Tylersmith - One of the best experts on this subject based on the ideXlab platform.

  • population structure stratification and introgression of human structural variation
    bioRxiv, 2020
    Co-Authors: Mohamed A Almarri, Anders Bergstrom, Javier Pradomartinez, Fengtang Yang, Alistair S Dunham, Yuan Chen, Matthew E Hurles, Chris Tylersmith
    Abstract:

    Abstract Structural variants contribute substantially to genetic Diversity and are important evolutionarily and medically, yet are still understudied. Here, we present a comprehensive analysis of deletions, duplications, insertions, inversions and non-reference unique insertions in the Human Genome Diversity Project (HGDP-CEPH) panel, a high-coverage dataset of 911 samples from 54 diverse worldwide populations. We identify in total 126,018 structural variants (25,588

  • population structure stratification and introgression of human structural variation in the hgdp
    bioRxiv, 2019
    Co-Authors: Mohamed A Almarri, Anders Bergstrom, Javier Pradomartinez, Alistair S Dunham, Yuan Chen, Chris Tylersmith, Yali Xue
    Abstract:

    Abstract Structural variants contribute substantially to genetic Diversity and are important evolutionarily and medically, yet are still understudied. Here, we present a comprehensive analysis of deletions, duplications, inversions and non-reference unique insertions in the Human Genome Diversity Project (HGDP-CEPH) panel, a high-coverage dataset of 910 samples from 54 diverse worldwide populations. We identify in total 61,801 structural variants, of which 61% are novel. Some reach high frequency and are private to continental groups or even individual populations, including a deletion in the maltase-glucoamylase gene MGAM, involved in starch digestion, in the South American Karitiana and a deletion in the Central African Mbuti in SIGLEC5, potentially increasing susceptibility to autoimmune diseases. We discover a dynamic range of copy number expansions and find cases of regionally-restricted runaway duplications, for example, 18 copies near the olfactory receptor OR7D2 in East Asia and in the clinically-relevant HCAR2 in Central Asia. We identify highly-stratified putatively introgressed variants from Neanderthals or Denisovans, some of which, like a deletion within AQR in Papuans, are almost fixed in individual populations. Finally, by de novo assembly of 25 genomes using linked-read sequencing we discover 1631 breakpoint-resolved unique insertions, in aggregate accounting for 1.9 Mb of sequence absent from the GRCh38 reference. These insertions show population structure and some reside in functional regions, illustrating the limitation of a single human reference and the need for high-quality genomes from diverse populations to fully discover and understand human genetic variation.

  • chromosome wide characterization of y str mutation rates
    bioRxiv, 2016
    Co-Authors: Thomas Willems, Melissa Gymrek, Chris Tylersmith, David G Poznik, Yaniv Erlich
    Abstract:

    Short Tandem Repeats (STRs) are mutation-prone loci that span nearly 1% of the human genome. Previous studies have estimated the mutation rates of highly polymorphic STRs using capillary electrophoresis and pedigree-based designs. While this work has provided insights into the mutational dynamics of highly mutable STRs, the mutation rates of most others remain unknown. Here, we harnessed whole-genome sequencing data to estimate the mutation rates of more than 4,500 Y-chromosome STRs (Y-STRs) with 2-6 base pair repeat units. To this end, we developed MUTEA, a new algorithm that infers STR mutation rates from population-scale high-throughput sequencing data using a high-resolution SNP-based phylogeny. After extensive intrinsic and extrinsic validations, we used MUTEA to estimate the mutation rates of STRs across the Y-chromosome using data from the 1000 Genomes Project and the Simons Genome Diversity Project. In total, we analyzed evolutionary data for over 222,000 meioses to yield the largest set of Y-STR mutation rate estimates to date. We found that the average mutation rate of polymorphic Y-STRs is an order of magnitude lower than estimates from prior studies. Using our ascertainment-free estimates, we identified determinants of STR mutation rates and built a model to predict rates for STRs across the genome. Our Projection indicates that the load of de novo STR mutations exceeds the load of all other known variants. We also identified new Y-STRs for forensics and genetic genealogy, assessed the ability to differentiate between the Y-chromosomes of father-son pairs, and imputed Y-STR genotypes.

  • chromosome wide characterization of y str mutation rates using ultra deep genealogies
    bioRxiv, 2016
    Co-Authors: Thomas Willems, Melissa Gymrek, Chris Tylersmith, David G Poznik, Yaniv Erlich
    Abstract:

    Although the utility of short tandem repeats on the Y-chromosome (Y-STRs) has long been recognized and leveraged in forensics, genealogy and paternity testing, the bulk of these applications have relied on only a few dozen loci identified as having remarkably high mutation rates. Recent efforts have expanded the set of Y-STRs with known mutation rates to two hundred markers, but the limited throughput of the capillary method for estimating mutation rates has left the mutability of most Y-STRs uncharacterized, particularly those with dinucleotide repeat units. To address this limitation, we developed a novel method capable of concurrently estimating the mutation rates of all Y-STRs by leveraging population-scale whole-genome sequencing data. Extensive simulations confirmed that our method robustly accounts for PCR stutter artifacts and obtains unbiased mutation rate estimates. Application of the method to orthogonal datasets from the 1000 Genomes Project and Simons Genome Diversity Project utilized evolutionary data from over 250,000 meioses to estimate the mutation rates of more than 700 Y-STRs with 2-6 base pair repeat units, yielding the largest such set to date. Comparison of these estimates with those from father-son studies indicated a high degree of concordance for loci that have been previously characterized. In addition, we identified nearly 100 previously uncharacterized Y-STRs with per-generation mutation rates greater than 1 in 3000. Altogether, our study provides a broadly applicable method for estimating Y-STR mutation rates from whole-genome sequencing cohorts, outlines a framework for imputing Y-STRs, vastly expands the number of identified loci with high discriminative power and provides the first chromosome-wide characterization of the mutation rates of dinucleotide short tandem repeats.

  • a worldwide survey of human male demographic history based on y snp and y str data from the hgdp ceph populations
    Molecular Biology and Evolution, 2010
    Co-Authors: Wentao Shi, Peter De Knijff, Manfred Kayser, Qasim Ayub, Mark Vermeulen, Rongguang Shao, Sofia Zuniga, Kristiaan J Van Der Gaag, Yali Xue, Chris Tylersmith
    Abstract:

    We have investigated human male demographic history using 590 males from 51 populations in the Human Genome Diversity Project - Centre d'Etude du Polymorphisme Humain worldwide panel, typed with 37 Y-chromosomal Single Nucleotide Polymorphisms and 65 Y-chromosomal Short Tandem Repeats and analyzed with the program Bayesian Analysis of Trees With Internal Node Generation. The general patterns we observe show a gradient from the oldest population time to the most recent common ancestors (TMRCAs) and expansion times together with the largest effective population sizes in Africa, to the youngest times and smallest effective population sizes in the Americas. These parameters are significantly negatively correlated with distance from East Africa, and the patterns are consistent with most other studies of human variation and history. In contrast, growth rate showed a weaker correlation in the opposite direction. Y-lineage Diversity and TMRCA also decrease with distance from East Africa, supporting a model of expansion with serial founder events starting from this source. A number of individual populations diverge from these general patterns, including previously documented examples such as recent expansions of the Yoruba in Africa, Basques in Europe, and Yakut in Northern Asia. However, some unexpected demographic histories were also found, including low growth rates in the Hazara and Kalash from Pakistan and recent expansion of the Mozabites in North Africa.