The Experts below are selected from a list of 29079 Experts worldwide ranked by ideXlab platform

Rounak Dey - One of the best experts on this subject based on the ideXlab platform.

  • an efficient and accurate frailty model approach for genome wide survival association analysis controlling for population structure and relatedness in large scale Biobanks
    bioRxiv, 2020
    Co-Authors: Rounak Dey, Wei Zhou, Tuomo Kiiskinen, Aki S Havulinna, Amanda F Elliott, Juha Karjalainen
    Abstract:

    With decades of electronic health records linked to genetic data, large Biobanks provide unprecedented opportunities for systematically understanding the genetics of the natural history of complex diseases. Genome-wide survival association analysis can identify genetic variants associated with ages of onset, disease progression and lifespan. We developed an efficient and accurate frailty (random effects) model approach for genome-wide survival association analysis of censored time-to-event (TTE) phenotypes in large Biobanks by accounting for both population structure and relatedness. Our method utilizes state-of-the-art optimization strategies to reduce the computational cost. The saddlepoint approximation is used to allow for analysis of heavily censored phenotypes (>90%) and low frequency variants (down to minor allele count 20). We demonstrated the performance of our method through extensive simulation studies and analysis of five TTE phenotypes, including lifespan, with heavy censoring rates (90.9% to 99.8%) on ~400,000 UK Biobank participants with white British ancestry and ~180,000 samples in FinnGen, respectively. We further performed genome-wide association analysis for 871 TTE phenotypes in UK Biobank and presented the genome-wide scale phenome-wide association (PheWAS) results with the PheWeb browser.

  • a fast and accurate method for genome wide scale phenome wide g e analysis and its application to uk Biobank
    American Journal of Human Genetics, 2019
    Co-Authors: Zhangchen Zhao, Rounak Dey, Lars G Fritsche, Bhramar Mukherjee, Seunggeun Lee
    Abstract:

    The etiology of most complex diseases involves genetic variants, environmental factors, and gene-environment interaction (G × E) effects. Compared with marginal genetic association studies, G × E analysis requires more samples and detailed measure of environmental exposures, and this limits the possible discoveries. Large-scale population-based Biobanks with detailed phenotypic and environmental information, such as UK-Biobank, can be ideal resources for identifying G × E effects. However, due to the large computation cost and the presence of case-control imbalance, existing methods often fail. Here we propose a scalable and accurate method, SPAGE (SaddlePoint Approximation implementation of G × E analysis), that is applicable for genome-wide scale phenome-wide G × E studies. SPAGE fits a genotype-independent logistic model only once across the genome-wide analysis in order to reduce computation cost, and SPAGE uses a saddlepoint approximation (SPA) to calibrate the test statistics for analysis of phenotypes with unbalanced case-control ratios. Simulation studies show that SPAGE is 33–79 times faster than the Wald test and 72–439 times faster than the Firth’s test, and SPAGE can control type I error rates at the genome-wide significance level even when case-control ratios are extremely unbalanced. Through the analysis of UK-Biobank data of 344,341 white British European-ancestry samples, we show that SPAGE can efficiently analyze large samples while controlling for unbalanced case-control ratios.

  • robust meta analysis of Biobank based genome wide association studies with unbalanced binary phenotypes
    Genetic Epidemiology, 2019
    Co-Authors: Rounak Dey, Wei Zhou, Lars G Fritsche, J B Nielsen, Huanhuan Zhu, Cristen J Willer, Seunggeun Lee
    Abstract:

    With the availability of large-scale Biobanks, genome-wide scale phenome-wide association studies are being instrumental in discovering novel genetic variants associated with clinical phenotypes. As increasing number of such association results from different Biobanks become available, methods to meta-analyse those association results is of great interest. Because the binary phenotypes in Biobank-based studies are mostly unbalanced in their case-control ratios, very few methods can provide well-calibrated tests for associations. For example, traditional Z-score-based meta-analysis often results in conservative or anticonservative Type I error rates in such unbalanced scenarios. We propose two meta-analysis strategies that can efficiently combine association results from Biobank-based studies with such unbalanced phenotypes, using the saddlepoint approximation-based score test method. Our first method involves sharing the overall genotype counts from each study, and the second method involves sharing an approximation of the distribution of the score test statistic from each study using cubic Hermite splines. We compare our proposed methods with a traditional Z-score-based meta-analysis strategy using numerical simulations and real data applications, and demonstrate the superior performance of our proposed methods in terms of Type I error control.

  • efficiently controlling for case control imbalance and sample relatedness in large scale genetic association studies
    Nature Genetics, 2018
    Co-Authors: Wei Zhou, Rounak Dey, Lars G Fritsche, Jonas B Nielsen, Maiken E Gabrielsen, Brooke N Wolford, Jonathon Lefaive, Peter Vandehaar, Sarah A Gagliano, Aliya Gifford
    Abstract:

    In genome-wide association studies (GWAS) for thousands of phenotypes in large Biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, the linear mixed model and the recently proposed logistic mixed model, perform poorly; they produce large type I error rates when used to analyze unbalanced case-control phenotypes. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approximation to calibrate the distribution of score test statistics. This method, SAIGE (Scalable and Accurate Implementation of GEneralized mixed model), provides accurate P values even when case-control ratios are extremely unbalanced. SAIGE uses state-of-art optimization strategies to reduce computational costs; hence, it is applicable to GWAS for thousands of phenotypes by large Biobanks. Through the analysis of UK Biobank data of 408,961 samples from white British participants with European ancestry for > 1,400 binary phenotypes, we show that SAIGE can efficiently analyze large sample data, controlling for unbalanced case-control ratios and sample relatedness.

  • efficiently controlling for case control imbalance and sample relatedness in large scale genetic association studies
    bioRxiv, 2017
    Co-Authors: Wei Zhou, Rounak Dey, Lars G Fritsche, Jonas B Nielsen, Maiken E Gabrielsen, Brooke N Wolford, Jonathon Lefaive, Peter Vandehaar, Aliya Gifford, Lisa Bastarache
    Abstract:

    In genome-wide association studies (GWAS) for thousands of phenotypes in large Biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, linear mixed model and the recently proposed logistic mixed model, perform poorly -- producing large type I error rates -- in the analysis of phenotypes with unbalanced case-control ratios. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approximation (SPA) to calibrate the distribution of score test statistics. This method, SAIGE, provides accurate p-values even when case-control ratios are extremely unbalanced. It utilizes state-of-art optimization strategies to reduce computational time and memory cost of generalized mixed model. The computation cost linearly depends on sample size, and hence can be applicable to GWAS for thousands of phenotypes by large Biobanks. Through the analysis of UK-Biobank data of 408,961 white British European-ancestry samples, we show that SAIGE can efficiently analyze large sample data, controlling for unbalanced case-control ratios and sample relatedness.

Seunggeun Lee - One of the best experts on this subject based on the ideXlab platform.

  • a fast and accurate method for genome wide scale phenome wide g e analysis and its application to uk Biobank
    American Journal of Human Genetics, 2019
    Co-Authors: Zhangchen Zhao, Rounak Dey, Lars G Fritsche, Bhramar Mukherjee, Seunggeun Lee
    Abstract:

    The etiology of most complex diseases involves genetic variants, environmental factors, and gene-environment interaction (G × E) effects. Compared with marginal genetic association studies, G × E analysis requires more samples and detailed measure of environmental exposures, and this limits the possible discoveries. Large-scale population-based Biobanks with detailed phenotypic and environmental information, such as UK-Biobank, can be ideal resources for identifying G × E effects. However, due to the large computation cost and the presence of case-control imbalance, existing methods often fail. Here we propose a scalable and accurate method, SPAGE (SaddlePoint Approximation implementation of G × E analysis), that is applicable for genome-wide scale phenome-wide G × E studies. SPAGE fits a genotype-independent logistic model only once across the genome-wide analysis in order to reduce computation cost, and SPAGE uses a saddlepoint approximation (SPA) to calibrate the test statistics for analysis of phenotypes with unbalanced case-control ratios. Simulation studies show that SPAGE is 33–79 times faster than the Wald test and 72–439 times faster than the Firth’s test, and SPAGE can control type I error rates at the genome-wide significance level even when case-control ratios are extremely unbalanced. Through the analysis of UK-Biobank data of 344,341 white British European-ancestry samples, we show that SPAGE can efficiently analyze large samples while controlling for unbalanced case-control ratios.

  • robust meta analysis of Biobank based genome wide association studies with unbalanced binary phenotypes
    Genetic Epidemiology, 2019
    Co-Authors: Rounak Dey, Wei Zhou, Lars G Fritsche, J B Nielsen, Huanhuan Zhu, Cristen J Willer, Seunggeun Lee
    Abstract:

    With the availability of large-scale Biobanks, genome-wide scale phenome-wide association studies are being instrumental in discovering novel genetic variants associated with clinical phenotypes. As increasing number of such association results from different Biobanks become available, methods to meta-analyse those association results is of great interest. Because the binary phenotypes in Biobank-based studies are mostly unbalanced in their case-control ratios, very few methods can provide well-calibrated tests for associations. For example, traditional Z-score-based meta-analysis often results in conservative or anticonservative Type I error rates in such unbalanced scenarios. We propose two meta-analysis strategies that can efficiently combine association results from Biobank-based studies with such unbalanced phenotypes, using the saddlepoint approximation-based score test method. Our first method involves sharing the overall genotype counts from each study, and the second method involves sharing an approximation of the distribution of the score test statistic from each study using cubic Hermite splines. We compare our proposed methods with a traditional Z-score-based meta-analysis strategy using numerical simulations and real data applications, and demonstrate the superior performance of our proposed methods in terms of Type I error control.

Saskia C Sanderson - One of the best experts on this subject based on the ideXlab platform.

  • public attitudes toward consent and data sharing in Biobank research a large multi site experimental survey in the us
    American Journal of Human Genetics, 2017
    Co-Authors: Ellen Wright Clayton, Saskia C Sanderson, Nathaniel D Mercaldo, Armand Matheny H Antommaria, Sharon Aufox, Murray H Brilliant, Diego Campos, David Carrell
    Abstract:

    Individuals participating in Biobanks and other large research projects are increasingly asked to provide broad consent for open-ended research use and widespread sharing of their biosamples and data. We assessed willingness to participate in a Biobank using different consent and data sharing models, hypothesizing that willingness would be higher under more restrictive scenarios. Perceived benefits, concerns, and information needs were also assessed. In this experimental survey, individuals from 11 US healthcare systems in the Electronic Medical Records and Genomics (eMERGE) Network were randomly allocated to one of three hypothetical scenarios: tiered consent and controlled data sharing; broad consent and controlled data sharing; or broad consent and open data sharing. Of 82,328 eligible individuals, exactly 13,000 (15.8%) completed the survey. Overall, 66% (95% CI: 63%–69%) of population-weighted respondents stated they would be willing to participate in a Biobank; willingness and attitudes did not differ between respondents in the three scenarios. Willingness to participate was associated with self-identified white race, higher educational attainment, lower religiosity, perceiving more research benefits, fewer concerns, and fewer information needs. Most (86%, CI: 84%–87%) participants would want to know what would happen if a researcher misused their health information; fewer (51%, CI: 47%–55%) would worry about their privacy. The concern that the use of broad consent and open data sharing could adversely affect participant recruitment is not supported by these findings. Addressing potential participants’ concerns and information needs and building trust and relationships with communities may increase acceptance of broad consent and wide data sharing in Biobank research.

Scott Y H Kim - One of the best experts on this subject based on the ideXlab platform.

  • understanding the public s reservations about broad consent and study by study consent for donations to a Biobank results of a national survey
    PLOS ONE, 2016
    Co-Authors: Raymond De Vries, Tom Tomlinson, Hyungjin Myra Kim, Chris Krenz, Diana K Haggerty, Kerry A Ryan, Scott Y H Kim
    Abstract:

    Researchers and policymakers do not agree about the most appropriate way to get consent for the use of donations to a Biobank. The most commonly used method is blanket—or broad—consent where donors allow their donation to be used for any future research approved by the Biobank. This approach does not account for the fact that some donors may have moral concerns about the uses of their biospecimens. This problem can be avoided using “real-time”—or study-by-study—consent, but this policy places a significant burden on Biobanks. In order to better understand the public’s preferences regarding Biobank consent policy, we surveyed a sample that was representative of the population of the United States. Respondents were presented with 5 Biobank consent policies and were asked to indicate which policies were acceptable/unacceptable and to identify the best/worst policies. They were also given 7 research scenarios that could create moral concern (e.g. research intending to make abortions safer and more effective) and asked how likely they would be to provide broad consent knowing that their donation might be used in that research. Substantial minorities found both broad and study-by-study consent to be unacceptable and identified those two options as the worst policies. Furthermore, while the type of moral concern (e.g., regarding abortion, the commercial use of donations, or stem cell research) had no effect on policy preferences, an increase in the number of research scenarios generating moral concerns was related to an increased likelihood of finding broad consent to be the worst policy. The rejection of these ethically problematic and costly extremes is good news for Biobanks. The challenge now is to design a policy that combines consent with access to information in a way that assures potential donors that their interests and moral concerns are being respected.

Umit Topaloglu - One of the best experts on this subject based on the ideXlab platform.

  • developing a semantically rich ontology for the Biobank administration domain
    Journal of Biomedical Semantics, 2013
    Co-Authors: Mathias Brochhausen, Roxana Merinomartinez, Loreana Norlin, Martin N Fransson, Nitin Kanaskar, Mikael Eriksson, Roger A Hall, Sanela Kjellqvist, Maria Hortlund, Umit Topaloglu
    Abstract:

    Background: Biobanks are a critical resource for translational science. Recently, semantic web technologies such as ontologies have been found useful in retrieving research data from Biobanks. However, recent research has also shown that there is a lack of data about the administrative aspects of Biobanks. These data would be helpful to answer research-relevant questions such as what is the scope of specimens collected in a Biobank, what is the curation status of the specimens, and what is the contact information for curators of Biobanks. Our use cases include giving researchers the ability to retrieve key administrative data (e.g. contact information, contact's affiliation, etc.) about the Biobanks where specific specimens of interest are stored. Thus, our goal is to provide an ontology that represents the administrative entities in Biobanking and their relations. We base our ontology development on a set of 53 data attributes called MIABIS, which were in part the result of semantic integration efforts of the European Biobanking and Biomolecular Resources Research Infrastructure (BBMRI). The previous work on MIABIS provided the domain analysis for our ontology. We report on a test of our ontology against competency questions that we derived from the initial BBMRI use cases. Future work includes additional ontology development to answer additional competency questions from these use cases. Results: We created an open-source ontology of Biobank administration called Ontologized MIABIS (OMIABIS) coded in OWL 2.0 and developed according to the principles of the OBO Foundry. It re-uses pre-existing ontologies when possible in cooperation with developers of other ontologies in related domains, such as the Ontology of Biomedical Investigation. OMIABIS provides a formalized representation of Biobanks and their administration. Using the ontology and a set of Description Logic queries derived from the competency questions that we identified, we were able to retrieve test data with perfect accuracy. In addition, we began development of a mapping from the ontology to pre-existing Biobank data structures commonly used in the U.S. Conclusions: In conclusion, we created OMIABIS, an ontology of Biobank administration. We found that basing its development on pre-existing resources to meet the BBMRI use cases resulted in a Biobanking ontology that is re-useable in environments other than BBMRI. Our ontology retrieved all true positives and no false positives when queried according to the competency questions we derived from the BBMRI use cases. Mapping OMIABIS to a data structure used for biospecimen collections in a medical center in Little Rock, AR showed adequate coverage of our ontology.

  • developing a semantically rich ontology for the Biobank administration domain
    Journal of Biomedical Semantics, 2013
    Co-Authors: Mathias Brochhausen, Roxana Merinomartinez, Loreana Norlin, Martin N Fransson, Nitin Kanaskar, Mikael Eriksson, Roger A Hall, Sanela Kjellqvist, Maria Hortlund, Umit Topaloglu
    Abstract:

    Background Biobanks are a critical resource for translational science. Recently, semantic web technologies such as ontologies have been found useful in retrieving research data from Biobanks. However, recent research has also shown that there is a lack of data about the administrative aspects of Biobanks. These data would be helpful to answer research-relevant questions such as what is the scope of specimens collected in a Biobank, what is the curation status of the specimens, and what is the contact information for curators of Biobanks. Our use cases include giving researchers the ability to retrieve key administrative data (e.g. contact information, contact's affiliation, etc.) about the Biobanks where specific specimens of interest are stored. Thus, our goal is to provide an ontology that represents the administrative entities in Biobanking and their relations. We base our ontology development on a set of 53 data attributes called MIABIS, which were in part the result of semantic integration efforts of the European Biobanking and Biomolecular Resources Research Infrastructure (BBMRI). The previous work on MIABIS provided the domain analysis for our ontology. We report on a test of our ontology against competency questions that we derived from the initial BBMRI use cases. Future work includes additional ontology development to answer additional competency questions from these use cases.