The Experts below are selected from a list of 10872 Experts worldwide ranked by ideXlab platform
Kristian B Filion - One of the best experts on this subject based on the ideXlab platform.
-
acetaminophen use during pregnancy and the risk of attention deficit hyperactivity disorder a causal association or bias
Paediatric and Perinatal Epidemiology, 2020Co-Authors: Kristian B Filion, Robert W Platt, Reem MasarwaAbstract:Background The association between acetaminophen use during pregnancy and the development of attention deficit hyperactivity disorder (ADHD) in the offspring may be due to bias. Objectives The primary objective was to assess the role of potential unmeasured confounding in the estimation of the association between acetaminophen use during pregnancy and the risk of ADHD, through bias analysis. The secondary objective was to assess the roles of selection bias and exposure misclassification. Data sources We searched MEDLINE, Embase, Scopus, and the Cochrane Library up to December 2018. Study selection and data extraction We included observational studies examining the association between acetaminophen use during pregnancy and the risk of ADHD. Synthesis We meta-analysed data across studies, using random-effects model. We conducted a bias analysis to studies that did not adjust for important Confounders, to explore systematic errors related to unmeasured confounding, selection bias, and exposure misclassification. Results The search resulted in seven studies included in our meta-analysis. When adjusted estimates were pooled across all studies, the risk ratio (RR) for ADHD was 1.35 (95% confidence interval [CI] 1.25, 1.46; I2 = 48%). Sensitivity analysis for unmeasured confounding in this meta-analysis showed that a Confounder of 1.69 on the RR scale would reduce to 10% the proportion of studies with a true effect size of RR >1.10. Unmeasured confounding bias analysis decreased the point estimate in five of the seven studies and increased in two studies, suggesting that the observed association could be confounded by parental ADHD. Unadjusted and bias-corrected risk ratios (bcRRs) were: RR = 1.34, bcRR = 1.13; RR = 1.51, bcRR = 1.17; RR = 1.63, bcRR = 1.38; RR = 1.44, bcRR = 1.17; RR = 1.16, bcRR = 1.18; RR = 1.25, bcRR = 1.05; and RR = 0.99, bcRR = 1.18. Conclusions Bias analysis suggests that the previously reported association between acetaminophen use during pregnancy and an increased risk of ADHD in the offspring may be due to unmeasured confounding. Our ability to conclude a causal association is limited.
-
visualization tool of variable selection in bias variance tradeoff for inverse probability weights
Annals of Epidemiology, 2020Co-Authors: Kristian B Filion, Lisa M Bodnar, Maria M Brooks, Robert W Platt, Katherine P Himes, Ashley I NaimiAbstract:Abstract Purpose Inversed probability weighted (IPW) estimators are commonly used to adjust for time-fixed or time-varying Confounders. However, in high-dimensional settings, including all identified Confounders may result in unstable weights leading to higher variance. We aimed to develop a visualization tool demonstrating the impact of each Confounder on the bias and variance of IPW estimates, as well as the propensity score overlap. Methods A SAS macro was developed for this visualization tool and we demonstrate how this tool can be used to identify potentially problematic Confounders of the association of statin use after myocardial infarction on one-year mortality in a plasmode simulation study using a cohort of 39,792 patients from the UK (1998–2012). Results Through the tool's output, we can identify problematic Confounders (two instrumental variables) and important Confounders by comparing the estimated psuedo MSE with that from the fully adjusted model and propensity score overlap plot. Conclusion Our results suggest that the analytic impact of all Confounders should be considered carefully when fitting IPW estimators.
-
multiple imputation for systematically missing Confounders within a distributed data drug safety network a simulation study and real world example
Pharmacoepidemiology and Drug Safety, 2020Co-Authors: Matthew H Secrest, Robert W Platt, Pauline Reynier, Colin R Dormuth, Andrea Benedetti, Kristian B FilionAbstract:Purpose In distributed data networks, some data sites may be systematically missing important Confounders that are captured by other sites in the network (eg, body mass index [BMI]). Multiple imputation may help repair bias in these scenarios. However, multiple imputation has not been described for distributed data networks where data access restrictions prevent centralized analysis. Methods We conducted a simulation study and a real-world analysis using the UK's Clinical Practice Research Datalink to evaluate multiple imputation for Confounders that are systematically missing from a subset of data sites in mock distributed data networks. The simulation study addressed univariate missing data, while the real-world analysis addressed multivariate missing data. Both studies were designed as retrospective cohort studies of the effect of current statin use on the risk of myocardial infarction among patients with newly treated type 2 diabetes. Results In our simulation study, multiple imputation repaired bias from missing BMI in all scenarios, with a median bias reduction of 118% in the default scenario. In our real-world study, the multiply imputed analysis (hazard ratio [HR]: 0.86; 95% confidence interval [CI], 0.69-1.08) was closer to the analysis that considered the true Confounder values (HR: 0.85; 95% CI, 0.66-1.10) than the analysis that ignored them (HR: 0.93; 95% CI, 0.73-1.20). Conclusions Multiple imputation adapted to distributed data settings is a feasible method to reduce bias from unmeasured but measurable Confounders when at least one database contains the variables of interest. Further research is needed to evaluate its validity in real distributed data networks.
Marvin Bertin - One of the best experts on this subject based on the ideXlab platform.
-
machine learning enables detection of early stage colorectal cancer by whole genome sequencing of plasma cell free dna
BMC Cancer, 2019Co-Authors: Nathan Wan, David E Weinberg, Tzuyu Liu, Katherine E Niehaus, Eric A Ariazi, Daniel Delubac, Ajay Kannan, Brandon White, Mitch Bailey, Marvin BertinAbstract:Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N = 546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and Confounder-based cross-validations to assess generalization performance. In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91–0.93) with a mean sensitivity of 85% (95% CI 83–86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these Confounders and provided a more accurate assessment of generalization performance. A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar Confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
-
machine learning enables detection of early stage colorectal cancer by whole genome sequencing of plasma cell free dna
bioRxiv, 2018Co-Authors: Nathan Wan, David E Weinberg, Tzuyu Liu, Katherine E Niehaus, Eric A Ariazi, Daniel Delubac, Ajay Kannan, Brandon White, Mitch Bailey, Marvin BertinAbstract:Background: Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Methods: Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N=546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and Confounder-based cross-validation to assess generalization performance. Results: In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91-0.93) with a mean sensitivity of 85% (95% CI 83-86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these Confounders and provided a more accurate assessment of generalization performance. Conclusions: A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar Confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
Gerda Claeskens - One of the best experts on this subject based on the ideXlab platform.
-
on model selection and model misspecification in causal inference
Statistical Methods in Medical Research, 2012Co-Authors: Stijn Vansteelandt, Maarten Bekaert, Gerda ClaeskensAbstract:Standard variable selection procedures, primarily developed for the construction of outcome prediction models, are routinely applied when assessing exposure effects in observational studies. We argue that this tradition is sub-optimal and prone to yield bias in exposure effect estimators as well as their corresponding uncertainty estimators. We weigh the pros and cons of Confounder-selection procedures and propose a procedure directly targeting the quality of the exposure effect estimator. We further demonstrate that certain strategies for inferring causal effects have the desirable features (a) of producing (approximately) valid confidence intervals, even when the Confounder-selection process is ignored, and (b) of being robust against certain forms of misspecification of the association of Confounders with both exposure and outcome.
-
on model selection and model misspecification in causal inference
Social Science Research Network, 2010Co-Authors: Stijn Vansteelandt, Maarten Bekaert, Gerda ClaeskensAbstract:Standard variable-selection procedures, primarily developed for the construction of outcome prediction models, are routinely applied when assessing exposure e®ects in observational studies. We argue that this tradition is sub-optimal and prone to yield bias in exposure effect estimates as well as their corresponding uncertainty estimates. We weigh the pros and cons of Confounder-selection procedures and propose a procedure directly targeting the quality of the exposure effect estimator. We further demonstrate that certain strategies for inferring causal effects have the desirable features (a) of producing (approximately) valid confidence intervals, even when the Confounder-selection process is ignored, and (b) of being robust against certain forms of misspecification of the association of Confounders with both exposure and outcome.
Nathan Wan - One of the best experts on this subject based on the ideXlab platform.
-
machine learning enables detection of early stage colorectal cancer by whole genome sequencing of plasma cell free dna
BMC Cancer, 2019Co-Authors: Nathan Wan, David E Weinberg, Tzuyu Liu, Katherine E Niehaus, Eric A Ariazi, Daniel Delubac, Ajay Kannan, Brandon White, Mitch Bailey, Marvin BertinAbstract:Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N = 546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and Confounder-based cross-validations to assess generalization performance. In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91–0.93) with a mean sensitivity of 85% (95% CI 83–86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these Confounders and provided a more accurate assessment of generalization performance. A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar Confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
-
machine learning enables detection of early stage colorectal cancer by whole genome sequencing of plasma cell free dna
bioRxiv, 2018Co-Authors: Nathan Wan, David E Weinberg, Tzuyu Liu, Katherine E Niehaus, Eric A Ariazi, Daniel Delubac, Ajay Kannan, Brandon White, Mitch Bailey, Marvin BertinAbstract:Background: Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Methods: Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N=546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and Confounder-based cross-validation to assess generalization performance. Results: In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91-0.93) with a mean sensitivity of 85% (95% CI 83-86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these Confounders and provided a more accurate assessment of generalization performance. Conclusions: A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar Confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
Simon G. Thompson - One of the best experts on this subject based on the ideXlab platform.
-
multiple imputation for handling systematically missing Confounders in meta analysis of individual participant data
Statistics in Medicine, 2013Co-Authors: Matthieu Rescherigon, Sanne A E Peters, Ian R. White, Jonathan W Bartlett, Simon G. ThompsonAbstract:A variable is ‘systematically missing’ if it is missing for all individuals within particular studies in an individual participant data meta-analysis. When a systematically missing variable is a potential Confounder in observational epidemiology, standard methods either fail to adjust the exposure–disease association for the potential Confounder or exclude studies where it is missing. We propose a new approach to adjust for systematically missing Confounders based on multiple imputation by chained equations. Systematically missing data are imputed via multilevel regression models that allow for heterogeneity between studies. A simulation study compares various choices of imputation model. An illustration is given using data from eight studies estimating the association between carotid intima media thickness and subsequent risk of cardiovascular events. Results are compared with standard methods and also with an extension of a published method that exploits the relationship between fully adjusted and partially adjusted estimated effects through a multivariate random effects meta-analysis model. We conclude that multiple imputation provides a practicable approach that can handle arbitrary patterns of systematic missingness. Bias is reduced by including sufficient between-study random effects in the imputation model. Copyright © 2013 John Wiley & Sons, Ltd.