The Experts below are selected from a list of 18171 Experts worldwide ranked by ideXlab platform
C J Battey - One of the best experts on this subject based on the ideXlab platform.
-
Minor Allele Frequency thresholds strongly affect population structure inference with genomic data sets
2019Co-Authors: Ethan Linck, C J BatteyAbstract:A common method of minimizing errors in large DNA sequence data sets is to drop variable sites with a Minor Allele Frequency (MAF) below some specified threshold. Although widespread, this procedure has the potential to alter downstream population genetic inferences and has received relatively little rigorous analysis. Here we use simulations and an empirical single nucleotide polymorphism data set to demonstrate the impacts of MAF thresholds on inference of population structure-often the first step in analysis of population genomic data. We find that model-based inference of population structure is confounded when singletons are included in the alignment, and that both model-based and multivariate analyses infer less distinct clusters when more stringent MAF cutoffs are applied. We propose that this behaviour is caused by the combination of a drop in the total size of the data matrix and by correlations between Allele frequencies and mutational age. We recommend a set of best practices for applying MAF filters in studies seeking to describe population structure with genomic data.
-
Minor Allele Frequency thresholds strongly affect population structure inference with genomic datasets
2017Co-Authors: Ethan Linck, C J BatteyAbstract:Across the genome, the effects of different evolutionary processes and historical events can result in different classes of genetic variants (or Alleles) characterized by their relative Frequency in a given population. As a result, population genetic inference can be strongly affected by biases in laboratory and bioinformatics treatments that affect the site Frequency spectrum, or SFS. Yet despite the widespread use of reduced-representation genomic datasets with nonmodel organisms, the potential consequences of these biases for downstream analyses remain poorly examined. Here, we assess the influence of Minor Allele Frequency (MAF) thresholds implemented during variant detection on inference of population structure. We use simulated and empirical datasets to evaluate the effect of MAF thresholds on the ability to discriminate among populations and quantify admixture with both model-based and non-model-based clustering methods. We find model-based inference of population structure is highly sensitive to choice of MAF, and may be confounded by either including singletons or excluding all rare Alleles. In contrast, non-model-based clustering is largely robust to MAF choice. Our results suggest that model-based inference of population structure can fail due to either natural demographic processes or assembly artifacts, with broad consequences for phylogeographic and population genetic studies using NGS data. We propose a simple hypothesis to explain this behavior and recommend a set of best practices for researchers seeking to describe population structure using reduced-representation libraries.
Richard A. Gibbs - One of the best experts on this subject based on the ideXlab platform.
-
whole genome sequence analysis of serum amino acid levels
2016Co-Authors: Bing Yu, Paul S De Vries, Ginger A Metcalf, Elena V Feofanova, Donna M Muzny, Lynne E Wagenknecht, Alanna C Morrison, Zhe Wang, Richard A. Gibbs, Eric BoerwinkleAbstract:Blood levels of amino acids are important biomarkers of disease and are influenced by synthesis, protein degradation, and gene–environment interactions. Whole genome sequence analysis of amino acid levels may establish a paradigm for analyzing quantitative risk factors. In a discovery cohort of 1872 African Americans and a replication cohort of 1552 European Americans we sequenced exons and whole genomes and measured serum levels of 70 amino acids. Rare and low-Frequency variants (Minor Allele Frequency ≤5%) were analyzed by three types of aggregating motifs defined by gene exons, regulatory regions, or genome-wide sliding windows. Common variants (Minor Allele Frequency >5%) were analyzed individually. Over all four analysis strategies, 14 gene–amino acid associations were identified and replicated. The 14 loci accounted for an average of 1.8% of the variance in amino acid levels, which ranged from 0.4 to 9.7%. Among the identified locus–amino acid pairs, four are novel and six have been reported to underlie known Mendelian conditions. These results suggest that there may be substantial genetic effects on amino acid levels in the general population that may underlie inborn errors of metabolism. We also identify a predicted promoter variant in AGA (the gene that encodes aspartylglucosaminidase) that is significantly associated with asparagine levels, with an effect that is independent of any observed coding variants. These data provide insights into genetic influences on circulating amino acid levels by integrating -omic technologies in a multi-ethnic population. The results also help establish a paradigm for whole genome sequence analysis of quantitative traits.
-
sequencing of 2 subclinical atherosclerosis candidate regions in 3669 individuals cohorts for heart and aging research in genomic epidemiology charge consortium targeted sequencing study
2014Co-Authors: Joshua C Bis, Donna M Muzny, Richard A. Gibbs, Charles C White, Nora Franceschini, Jennifer A Brody, Xiaoling Zhang, Jireh Santibanez, Xiaoming LiuAbstract:Background—Atherosclerosis, the precursor to coronary heart disease and stroke, is characterized by an accumulation of fatty cells in the arterial intimal-medial layers. Common carotid intima media thickness (cIMT) and plaque are subclinical atherosclerosis measures that predict cardiovascular disease events. Previously, genome-wide association studies demonstrated evidence for association with cIMT (SLC17A4) and plaque (PIK3CG). Methods and Results—We sequenced 120 kb around SLC17A4 (6p22.2) and 251 kb around PIK3CG (7q22.3) among 3669 European ancestry participants from the Atherosclerosis Risk in Communities (ARIC) study, Cardiovascular Health Study (CHS), and Framingham Heart Study (FHS) in Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium. Primary analyses focused on 438 common variants (Minor Allele Frequency ≥1%), which were independently meta-analyzed. A 3ʹ untranslated region CCDC71L variant (rs2286149), upstream from PIK3CG, was the most significant finding in cIMT (P=0.00033) and plaque (P=0.0004) analyses. A SLC17A4 intronic variant was also associated with cIMT (P=0.008). Both were in low linkage disequilibrium with the genome-wide association study single nucleotide polymorphisms. Gene-based tests including T1 count and sequence kernel association test for rare variants (Minor Allele Frequency <1%) did not yield statistically significant associations. However, we observed nominal associations for rare variants in CCDC71L and SLC17A3 with cIMT and of the entire 7q22 region with plaque (P=0.05). Conclusions—Common and rare variants in PIK3CG and SLC17A4 regions demonstrated modest association with subclinical atherosclerosis traits. Although not conclusive, these findings may help to understand the genetic architecture of regions previously implicated by genome-wide association studies and identify variants within these regions for further investigation in larger samples. (Circ Cardiovasc Genet. 2014;7:359-364.)
-
sequence variation in tmem18 in association with body mass index cohorts for heart and aging research in genomic epidemiology charge consortium targeted sequencing study
2014Co-Authors: Chingti Liu, Donna M Muzny, Alanna C Morrison, Jennifer A Brody, Kristin L Young, Matthias Olden, Mary K Wojczynski, Nancy L Heardcosta, Richard A. GibbsAbstract:Background— Genome-wide association studies for body mass index (BMI) previously identified a locus near TMEM18. We conducted targeted sequencing of this region to investigate the role of common, low-Frequency, and rare variants influencing BMI. Methods and Results— We sequenced TMEM18 and regions downstream of TMEM18 on chromosome 2 in 3976 individuals of European ancestry from 3 community-based cohorts (Atherosclerosis Risk in Communities, Cardiovascular Health Study, and Framingham Heart Study), including 200 adults selected for high BMI. We examined the association between BMI and variants identified in the region from nucleotide position 586 432 to 677 539 (hg18). Rare variants (Minor Allele Frequency, <1%) were analyzed using a burden test and the sequence kernel association test. Results from the 3 cohort studies were meta-analyzed. We estimate that mean BMI is 0.43 kg/m2 higher for each copy of the G Allele of single-nucleotide polymorphism rs7596758 (Minor Allele Frequency, 29%; P =3.46×10−4) using a Bonferroni threshold of P <4.6×10−4. Analyses conditional on previous genome-wide association study single-nucleotide polymorphisms associated with BMI in the region led to attenuation of this signal and uncovered another independent ( r 2<0.2), statistically significant association, rs186019316 ( P =2.11×10−4). Both rs186019316 and rs7596758 or proxies are located in transcription factor binding regions. No significant association with rare variants was found in either the exons of TMEM18 or the 3′ genome-wide association study region. Conclusions— Targeted sequencing around TMEM18 identified 2 novel BMI variants with possible regulatory function.
Donna M Muzny - One of the best experts on this subject based on the ideXlab platform.
-
whole genome sequence analysis of serum amino acid levels
2016Co-Authors: Bing Yu, Paul S De Vries, Ginger A Metcalf, Elena V Feofanova, Donna M Muzny, Lynne E Wagenknecht, Alanna C Morrison, Zhe Wang, Richard A. Gibbs, Eric BoerwinkleAbstract:Blood levels of amino acids are important biomarkers of disease and are influenced by synthesis, protein degradation, and gene–environment interactions. Whole genome sequence analysis of amino acid levels may establish a paradigm for analyzing quantitative risk factors. In a discovery cohort of 1872 African Americans and a replication cohort of 1552 European Americans we sequenced exons and whole genomes and measured serum levels of 70 amino acids. Rare and low-Frequency variants (Minor Allele Frequency ≤5%) were analyzed by three types of aggregating motifs defined by gene exons, regulatory regions, or genome-wide sliding windows. Common variants (Minor Allele Frequency >5%) were analyzed individually. Over all four analysis strategies, 14 gene–amino acid associations were identified and replicated. The 14 loci accounted for an average of 1.8% of the variance in amino acid levels, which ranged from 0.4 to 9.7%. Among the identified locus–amino acid pairs, four are novel and six have been reported to underlie known Mendelian conditions. These results suggest that there may be substantial genetic effects on amino acid levels in the general population that may underlie inborn errors of metabolism. We also identify a predicted promoter variant in AGA (the gene that encodes aspartylglucosaminidase) that is significantly associated with asparagine levels, with an effect that is independent of any observed coding variants. These data provide insights into genetic influences on circulating amino acid levels by integrating -omic technologies in a multi-ethnic population. The results also help establish a paradigm for whole genome sequence analysis of quantitative traits.
-
sequencing of 2 subclinical atherosclerosis candidate regions in 3669 individuals cohorts for heart and aging research in genomic epidemiology charge consortium targeted sequencing study
2014Co-Authors: Joshua C Bis, Donna M Muzny, Richard A. Gibbs, Charles C White, Nora Franceschini, Jennifer A Brody, Xiaoling Zhang, Jireh Santibanez, Xiaoming LiuAbstract:Background—Atherosclerosis, the precursor to coronary heart disease and stroke, is characterized by an accumulation of fatty cells in the arterial intimal-medial layers. Common carotid intima media thickness (cIMT) and plaque are subclinical atherosclerosis measures that predict cardiovascular disease events. Previously, genome-wide association studies demonstrated evidence for association with cIMT (SLC17A4) and plaque (PIK3CG). Methods and Results—We sequenced 120 kb around SLC17A4 (6p22.2) and 251 kb around PIK3CG (7q22.3) among 3669 European ancestry participants from the Atherosclerosis Risk in Communities (ARIC) study, Cardiovascular Health Study (CHS), and Framingham Heart Study (FHS) in Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium. Primary analyses focused on 438 common variants (Minor Allele Frequency ≥1%), which were independently meta-analyzed. A 3ʹ untranslated region CCDC71L variant (rs2286149), upstream from PIK3CG, was the most significant finding in cIMT (P=0.00033) and plaque (P=0.0004) analyses. A SLC17A4 intronic variant was also associated with cIMT (P=0.008). Both were in low linkage disequilibrium with the genome-wide association study single nucleotide polymorphisms. Gene-based tests including T1 count and sequence kernel association test for rare variants (Minor Allele Frequency <1%) did not yield statistically significant associations. However, we observed nominal associations for rare variants in CCDC71L and SLC17A3 with cIMT and of the entire 7q22 region with plaque (P=0.05). Conclusions—Common and rare variants in PIK3CG and SLC17A4 regions demonstrated modest association with subclinical atherosclerosis traits. Although not conclusive, these findings may help to understand the genetic architecture of regions previously implicated by genome-wide association studies and identify variants within these regions for further investigation in larger samples. (Circ Cardiovasc Genet. 2014;7:359-364.)
-
sequence variation in tmem18 in association with body mass index cohorts for heart and aging research in genomic epidemiology charge consortium targeted sequencing study
2014Co-Authors: Chingti Liu, Donna M Muzny, Alanna C Morrison, Jennifer A Brody, Kristin L Young, Matthias Olden, Mary K Wojczynski, Nancy L Heardcosta, Richard A. GibbsAbstract:Background— Genome-wide association studies for body mass index (BMI) previously identified a locus near TMEM18. We conducted targeted sequencing of this region to investigate the role of common, low-Frequency, and rare variants influencing BMI. Methods and Results— We sequenced TMEM18 and regions downstream of TMEM18 on chromosome 2 in 3976 individuals of European ancestry from 3 community-based cohorts (Atherosclerosis Risk in Communities, Cardiovascular Health Study, and Framingham Heart Study), including 200 adults selected for high BMI. We examined the association between BMI and variants identified in the region from nucleotide position 586 432 to 677 539 (hg18). Rare variants (Minor Allele Frequency, <1%) were analyzed using a burden test and the sequence kernel association test. Results from the 3 cohort studies were meta-analyzed. We estimate that mean BMI is 0.43 kg/m2 higher for each copy of the G Allele of single-nucleotide polymorphism rs7596758 (Minor Allele Frequency, 29%; P =3.46×10−4) using a Bonferroni threshold of P <4.6×10−4. Analyses conditional on previous genome-wide association study single-nucleotide polymorphisms associated with BMI in the region led to attenuation of this signal and uncovered another independent ( r 2<0.2), statistically significant association, rs186019316 ( P =2.11×10−4). Both rs186019316 and rs7596758 or proxies are located in transcription factor binding regions. No significant association with rare variants was found in either the exons of TMEM18 or the 3′ genome-wide association study region. Conclusions— Targeted sequencing around TMEM18 identified 2 novel BMI variants with possible regulatory function.
Ethan Linck - One of the best experts on this subject based on the ideXlab platform.
-
Minor Allele Frequency thresholds strongly affect population structure inference with genomic data sets
2019Co-Authors: Ethan Linck, C J BatteyAbstract:A common method of minimizing errors in large DNA sequence data sets is to drop variable sites with a Minor Allele Frequency (MAF) below some specified threshold. Although widespread, this procedure has the potential to alter downstream population genetic inferences and has received relatively little rigorous analysis. Here we use simulations and an empirical single nucleotide polymorphism data set to demonstrate the impacts of MAF thresholds on inference of population structure-often the first step in analysis of population genomic data. We find that model-based inference of population structure is confounded when singletons are included in the alignment, and that both model-based and multivariate analyses infer less distinct clusters when more stringent MAF cutoffs are applied. We propose that this behaviour is caused by the combination of a drop in the total size of the data matrix and by correlations between Allele frequencies and mutational age. We recommend a set of best practices for applying MAF filters in studies seeking to describe population structure with genomic data.
-
Minor Allele Frequency thresholds strongly affect population structure inference with genomic datasets
2017Co-Authors: Ethan Linck, C J BatteyAbstract:Across the genome, the effects of different evolutionary processes and historical events can result in different classes of genetic variants (or Alleles) characterized by their relative Frequency in a given population. As a result, population genetic inference can be strongly affected by biases in laboratory and bioinformatics treatments that affect the site Frequency spectrum, or SFS. Yet despite the widespread use of reduced-representation genomic datasets with nonmodel organisms, the potential consequences of these biases for downstream analyses remain poorly examined. Here, we assess the influence of Minor Allele Frequency (MAF) thresholds implemented during variant detection on inference of population structure. We use simulated and empirical datasets to evaluate the effect of MAF thresholds on the ability to discriminate among populations and quantify admixture with both model-based and non-model-based clustering methods. We find model-based inference of population structure is highly sensitive to choice of MAF, and may be confounded by either including singletons or excluding all rare Alleles. In contrast, non-model-based clustering is largely robust to MAF choice. Our results suggest that model-based inference of population structure can fail due to either natural demographic processes or assembly artifacts, with broad consequences for phylogeographic and population genetic studies using NGS data. We propose a simple hypothesis to explain this behavior and recommend a set of best practices for researchers seeking to describe population structure using reduced-representation libraries.
Xiaoming Liu - One of the best experts on this subject based on the ideXlab platform.
-
sequencing of 2 subclinical atherosclerosis candidate regions in 3669 individuals cohorts for heart and aging research in genomic epidemiology charge consortium targeted sequencing study
2014Co-Authors: Joshua C Bis, Donna M Muzny, Richard A. Gibbs, Charles C White, Nora Franceschini, Jennifer A Brody, Xiaoling Zhang, Jireh Santibanez, Xiaoming LiuAbstract:Background—Atherosclerosis, the precursor to coronary heart disease and stroke, is characterized by an accumulation of fatty cells in the arterial intimal-medial layers. Common carotid intima media thickness (cIMT) and plaque are subclinical atherosclerosis measures that predict cardiovascular disease events. Previously, genome-wide association studies demonstrated evidence for association with cIMT (SLC17A4) and plaque (PIK3CG). Methods and Results—We sequenced 120 kb around SLC17A4 (6p22.2) and 251 kb around PIK3CG (7q22.3) among 3669 European ancestry participants from the Atherosclerosis Risk in Communities (ARIC) study, Cardiovascular Health Study (CHS), and Framingham Heart Study (FHS) in Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium. Primary analyses focused on 438 common variants (Minor Allele Frequency ≥1%), which were independently meta-analyzed. A 3ʹ untranslated region CCDC71L variant (rs2286149), upstream from PIK3CG, was the most significant finding in cIMT (P=0.00033) and plaque (P=0.0004) analyses. A SLC17A4 intronic variant was also associated with cIMT (P=0.008). Both were in low linkage disequilibrium with the genome-wide association study single nucleotide polymorphisms. Gene-based tests including T1 count and sequence kernel association test for rare variants (Minor Allele Frequency <1%) did not yield statistically significant associations. However, we observed nominal associations for rare variants in CCDC71L and SLC17A3 with cIMT and of the entire 7q22 region with plaque (P=0.05). Conclusions—Common and rare variants in PIK3CG and SLC17A4 regions demonstrated modest association with subclinical atherosclerosis traits. Although not conclusive, these findings may help to understand the genetic architecture of regions previously implicated by genome-wide association studies and identify variants within these regions for further investigation in larger samples. (Circ Cardiovasc Genet. 2014;7:359-364.)