The Experts below are selected from a list of 24699 Experts worldwide ranked by ideXlab platform
Peter C. Austin - One of the best experts on this subject based on the ideXlab platform.
-
the use of Bootstrapping when using propensity score matching without replacement a simulation study
2014Co-Authors: Peter C. Austin, Dylan S SmallAbstract:Propensity-score matching is frequently used to estimate the effect of treatments, exposures, and interventions when using observational data. An important issue when using propensity-score matching is how to estimate the standard error of the estimated treatment effect. Accurate variance estimation permits construction of confidence intervals that have the advertised coverage rates and tests of statistical significance that have the correct type I error rates. There is disagreement in the literature as to how standard errors should be estimated. The Bootstrap is a commonly used resampling method that permits estimation of the sampling variability of estimated parameters. Bootstrap methods are rarely used in conjunction with propensity-score matching. We propose two different Bootstrap methods for use when using propensity-score matching without replacementand examined their performance with a series of Monte Carlo simulations. The first method involved drawing Bootstrap Samples from the matched pairs in the propensity-score-matched Sample. The second method involved drawing Bootstrap Samples from the original Sample and estimating the propensity score separately in each Bootstrap Sample and creating a matched Sample within each of these Bootstrap Samples. The former approach was found to result in estimates of the standard error that were closer to the empirical standard deviation of the sampling distribution of estimated effects.
-
Bootstrap model selection had similar performance for selecting authentic and noise variables compared to backward variable elimination a simulation study
2008Co-Authors: Peter C. AustinAbstract:Abstract Objective Researchers have proposed using Bootstrap resampling in conjunction with automated variable selection methods to identify predictors of an outcome and to develop parsimonious regression models. Using this method, multiple Bootstrap Samples are drawn from the original data set. Traditional backward variable elimination is used in each Bootstrap Sample, and the proportion of Bootstrap Samples in which each candidate variable is identified as an independent predictor of the outcome is determined. The performance of this method for identifying predictor variables has not been examined. Study Design and Setting Monte Carlo simulation methods were used to determine the ability of Bootstrap model selection methods to correctly identify predictors of an outcome when those variables that are selected for inclusion in at least 50% of the Bootstrap Samples are included in the final regression model. We compared the performance of the Bootstrap model selection method to that of conventional backward variable elimination. Results Bootstrap model selection tended to result in an approximately equal proportion of selected models being equal to the true regression model compared with the use of conventional backward variable elimination. Conclusion Bootstrap model selection performed comparatively to backward variable elimination for identifying the true predictors of a binary outcome.
-
using the Bootstrap to improve estimation and confidence intervals for regression coefficients selected using backwards variable elimination
2008Co-Authors: Peter C. AustinAbstract:Applied researchers frequently use automated model selection methods, such as backwards variable elimination, to develop parsimonious regression models. Statisticians have criticized the use of these methods for several reasons, amongst them are the facts that the estimated regression coefficients are biased and that the derived confidence intervals do not have the advertised coverage rates. We developed a method to improve estimation of regression coefficients and confidence intervals which employs backwards variable elimination in multiple Bootstrap Samples. In a given Bootstrap Sample, predictor variables that are not selected for inclusion in the final regression model have their regression coefficient set to zero. Regression coefficients are averaged across the Bootstrap Samples, and non-parametric percentile Bootstrap confidence intervals are then constructed for each regression coefficient. We conducted a series of Monte Carlo simulations to examine the performance of this method for estimating regression coefficients and constructing confidence intervals for variables selected using backwards variable elimination. We demonstrated that this method results in confidence intervals with superior coverage compared with those developed from conventional backwards variable elimination. We illustrate the utility of our method by applying it to a large Sample of subjects hospitalized with a heart attack.
-
automated variable selection methods for logistic regression produced unstable models for predicting acute myocardial infarction mortality
2004Co-Authors: Peter C. AustinAbstract:Abstract Objectives Automated variable selection methods are frequently used to determine the independent predictors of an outcome. The objective of this study was to determine the reproducibility of logistic regression models developed using automated variable selection methods. Study design and setting An initial set of 29 candidate variables were considered for predicting mortality after acute myocardial infarction (AMI). We drew 1,000 Bootstrap Samples from a dataset consisting of 4,911 patients admitted to hospital with an AMI. Using each Bootstrap Sample, logistic regression models predicting 30-day mortality were obtained using backward elimination, forward selection, and stepwise selection. The agreement between the different model selection methods and the agreement across the 1,000 Bootstrap Samples were compared. Results Using 1,000 Bootstrap Samples, backward elimination identified 940 unique models for predicting mortality. Similar results were obtained for forward and stepwise selection. Three variables were identified as independent predictors of mortality among all Bootstrap Samples. Over half the candidate prognostic variables were identified as independent predictors in less than half of the Bootstrap Samples. Conclusion Automated variable selection methods result in models that are unstable and not reproducible. The variables selected as independent predictors are sensitive to random fluctuations in the data.
Stephen M S Lee - One of the best experts on this subject based on the ideXlab platform.
-
stochastically optimal Bootstrap Sample size for shrinkage type statistics
2016Co-Authors: Bei Wei, Stephen M S LeeAbstract:In nonregular problems where the conventional $$n$$n out of $$n$$n Bootstrap is inconsistent, the $$m$$m out of $$n$$n Bootstrap provides a useful remedy to restore consistency. Conventionally, optimal choice of the Bootstrap Sample size $$m$$m is taken to be the minimiser of a frequentist error measure, estimation of which has posed a major difficulty hindering practical application of the $$m$$m out of $$n$$n Bootstrap method. Relatively little attention has been paid to a stronger, stochastic, version of the optimal Bootstrap Sample size, defined as the minimiser of an error measure calculated directly from the observed Sample. Motivated by this stronger notion of optimality, we develop procedures for calculating the stochastically optimal value of $$m$$m. Our procedures are shown to work under special forms of Edgeworth-type expansions which are typically satisfied by statistics of the shrinkage type. Theoretical and empirical properties of our methods are illustrated with three examples, namely the James---Stein estimator, the ridge regression estimator and the post-model-selection regression estimator.
-
variance estimation for Sample quantiles using the m out of n Bootstrap
2005Co-Authors: K Y Cheung, Stephen M S LeeAbstract:We consider the problem of estimating the variance of a Sample quantile calculated from a random Sample of sizen. Ther-th-order kernel-smoothed Bootstrap estimator is known to yield an impressively small relative error of orderO(n −r/(2r+1) ). It nevertheless requires strong smoothness conditions on the underlying density function, and has a performance very sensitive to the precise choice of the bandwidth. The unsmoothed Bootstrap has a poorer relative error of orderO(n −1/4), but works for less smooth density functions. We investigate a modified form of the Bootstrap, known as them out ofn Bootstrap, and show that it yields a relative error of order smaller thanO(n −1/4) under the same smoothness conditions required by the conventional unsmoothed Bootstrap on the density function, provided that the Bootstrap Sample sizem is of an appropriate order. The estimator permits exact, simulation-free, computation and has accuracy fairly insensitive to the precise choice ofm. A simulation study is reported to provide empirical comparison of the various methods.
Stephen Lee - One of the best experts on this subject based on the ideXlab platform.
-
optimal Bootstrap Sample size in construction of percentile confidence bounds
2001Co-Authors: Kamhin Chung, Stephen LeeAbstract:In traditional Bootstrap applications the size of a Bootstrap Sample equals the parent Sample size, n say. Recent studies have shown that using a Bootstrap Sample size different from n may sometimes provide a more satisfactory solution. In this paper we apply the latter approach to correct for coverage error in construction of Bootstrap confidence bounds. We show that the coverage error of a Bootstrap percentile method confidence bound, which is of order O(n−2/2) typically, can be reduced to O(n−1) by use of an optimal Bootstrap Sample size. A simulation study is conducted to illustrate our findings, which also suggest that the new method yields intervals of shorter length and greater stability compared to competitors of similar coverage accuracy.
Kamhin Chung - One of the best experts on this subject based on the ideXlab platform.
-
optimal Bootstrap Sample size in construction of percentile confidence bounds
2001Co-Authors: Kamhin Chung, Stephen LeeAbstract:In traditional Bootstrap applications the size of a Bootstrap Sample equals the parent Sample size, n say. Recent studies have shown that using a Bootstrap Sample size different from n may sometimes provide a more satisfactory solution. In this paper we apply the latter approach to correct for coverage error in construction of Bootstrap confidence bounds. We show that the coverage error of a Bootstrap percentile method confidence bound, which is of order O(n−2/2) typically, can be reduced to O(n−1) by use of an optimal Bootstrap Sample size. A simulation study is conducted to illustrate our findings, which also suggest that the new method yields intervals of shorter length and greater stability compared to competitors of similar coverage accuracy.
P K Mwanakatwe - One of the best experts on this subject based on the ideXlab platform.
-
the comparison study of the model selection criteria on the tobit regression model based on the Bootstrap Sample augmentation mechanisms
2021Co-Authors: P K MwanakatweAbstract:The statistical regression technique is an essential data fitting tool to explore the generation mechanism of the random phenomenon. Therefore, the model selection technique is becoming important. ...
-
model selection criteria of the standard censored regression model based on the Bootstrap Sample augmentation mechanism
2020Co-Authors: P K MwanakatweAbstract:The statistical regression technique is an extraordinarily essential data fitting tool to explore the potential possible generation mechanism of the random phenomenon. Therefore, the model selection or the variable selection is becoming extremely important so as to identify the most appropriate model with the most optimal explanation effect on the interesting response. In this paper, we discuss and compare the Bootstrap-based model selection criteria on the standard censored regression model (Tobit regression model) under the circumstance of limited observation information. The Monte Carlo numerical evidence demonstrates that the performances of the model selection criteria based on the Bootstrap Sample augmentation strategy will become more competitive than their alternative ones, such as the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC) etc. under the circumstance of the inadequate observation information. Meanwhile, the numerical simulation experiments further demonstrate that the model identification risk due to the deficiency of the data information, such as the high censoring rate and rather limited number of observations, can be adequately compensated by increasing the scientific computation cost in terms of the Bootstrap Sample augmentation strategies. We also apply the recommended Bootstrap-based model selection criterion on the Tobit regression model to fit the real fidelity dataset.