The Experts below are selected from a list of 106500 Experts worldwide ranked by ideXlab platform
John W Graham - One of the best experts on this subject based on the ideXlab platform.
-
Missing Data Analysis and design
2012Co-Authors: John W GrahamAbstract:An introduction to Analysis. - Planned Missing Data designs. -Practical issues and solutions for Analysis and design.
-
Missing Data Theory
Missing Data, 2012Co-Authors: John W GrahamAbstract:In this first chapter, I accomplish several goals. First, building on my 20+ years of work on Missing Data Analysis, I outline a nomenclature or system for talking about the theory underlying the modern Analysis of Missing Data. I intend for this nomenclature to be in plain English, but nevertheless to be an accurate representation of statistical theory relating to Missing Data Analysis. Second, I describe many of the main components of Missing Data theory, including the causes or mechanisms of Missingness. Two general methods for handling Missing Data, in particular multiple imputation (MI) and maximum-likelihood (ML) methods, have developed out of the Missing Data theory I describe here. And as will be clear from reading this book, I fully endorse these methods. For the remainder of this chapter, I challenge some of the commonly held beliefs relating to Missing Data theory and Missing Data Analysis, and make a case that the MI and ML procedures, which have started to become mainstream in statistical Analysis with Missing Data, are applicable in a much larger range of contexts that typically believed.
-
Missing Data Analysis making it work in the real world
Annual Review of Psychology, 2009Co-Authors: John W GrahamAbstract:This review presents a practical summary of the Missing Data literature, including a sketch of Missing Data theory and descriptions of normalmodel multiple imputation (MI) and maximum likelihood methods. Practical Missing Data Analysis issues are discussed, most notably the inclusion of auxiliary variables for improving power and reducing bias. Solutions are given for Missing Data challenges such as handling longitudinal, categorical, and clustered Data with normal-model MI; including interactions in the Missing Data model; and handling large numbers of variables. The discussion of attrition and nonignorable Missingness emphasizes the need for longitudinal diagnostics and for reducing the uncertainty about the Missing Data mechanism under attrition. Strategies suggested for reducing attrition bias include using auxiliary variables, collecting follow-up Data on a sample of those initially Missing, and collecting Data on intent to drop out. Suggestions are given for moving forward with research on Missing Data and attrition.
-
How Many Imputations are Really Needed? Some Practical Clarifications of Multiple Imputation Theory
Prevention Science, 2007Co-Authors: John W Graham, Allison E. Olchowski, Tamika D. GilreathAbstract:Multiple imputation (MI) and full information maximum likelihood (FIML) are the two most common approaches to Missing Data Analysis. In theory, MI and FIML are equivalent when identical models are tested using the same variables, and when m , the number of imputations performed with MI, approaches infinity. However, it is important to know how many imputations are necessary before MI and FIML are sufficiently equivalent in ways that are important to prevention scientists. MI theory suggests that small values of m , even on the order of three to five imputations, yield excellent results. Previous guidelines for sufficient m are based on relative efficiency, which involves the fraction of Missing information ( γ ) for the parameter being estimated, and m . In the present study, we used a Monte Carlo simulation to test MI models across several scenarios in which γ and m were varied. Standard errors and p-values for the regression coefficient of interest varied as a function of m , but not at the same rate as relative efficiency. Most importantly, statistical power for small effect sizes diminished as m became smaller, and the rate of this power falloff was much greater than predicted by changes in relative efficiency. Based our findings, we recommend that researchers using MI should perform many more imputations than previously considered sufficient. These recommendations are based on γ , and take into consideration one’s tolerance for a preventable power falloff (compared to FIML) due to using too few imputations.
-
Analysis With Missing Data in Prevention Research
NIDA research monograph, 1994Co-Authors: John W Graham, Scott M. Hofer, Andrea M. PiccininAbstract:Missing Data problems have been a thorn in the side of prevention researchers for years. Although some solutions for these problems have been available in the statistical literature, these solutions have not found their way into mainstream prevention research. This chapter is meant to serve as an introduction to the systematic application of the Missing Data Analysis solutions presented recently by Little and Rubin (1987) and others. The chapter does not describe a complete strategy, but it is relevant for (1) Missing Data Analysis with continuous (but not categorical) Data, (2) Data that are reasonably normally distributed, and (3) solutions for Missing Data problems for analyses related to the general linear model in particular, analyses that use (or can use) a covariance matrix as input. The examples in the chapter come from drug prevention research. The chapter discusses (1) the problem of wanting to ask respondents more questions than most individuals can answer; (2) the problem of attrition and some solutions; and (3) the problem of special measurement procedures that are too expensive or time consuming to obtain for all subjects. The authors end with several conclusions: Whenever possible, researchers should use the Expectation-Maximization (EM) algorithm (or other maximum likelihood procedure, including the multiple-group structural equation-modeling procedure or, where appropriate, multiple imputation, for analyses involving Missing Data [the chapter provides concrete examples]); If researchers must use other analyses, they should keep in mind that these others produce biased results and should not be relied upon for final analyses; When Data are Missing, the appropriate Missing Data Analysis procedures do not generate something out of nothing but do make the most out of the Data available; When Data are Missing, researchers should work hard (especially when planning a study) to find the cause of Missingness and include the cause in the Analysis models; and Researchers should sample the cases originally Missing (whenever possible) and adjust EM algorithm parameter estimates accordingly.
Hua-liang Wei - One of the best experts on this subject based on the ideXlab platform.
-
CSE/EUC/DCABES - Handling Missing Data in Multivariate Time Series Using a Vector Autoregressive Model Based Imputation (VAR-IM) Algorithm. Part II: VAR-IM Algorithm Versus Modern Methods
2016 IEEE Intl Conference on Computational Science and Engineering (CSE) and IEEE Intl Conference on Embedded and Ubiquitous Computing (EUC) and 15th , 2016Co-Authors: Faraj Bashir, Hua-liang WeiAbstract:This part of the paper introduces a comparison of VAR-IM algorithm with modern techniques used for Missing Data Analysis. Quantitative methods are usually developed based on some fundamental understanding of the statistical Analysis of the Missing Data. Various types of quantitative methods such as K nearest neighbour (KNN), (Multivariate Autoregressive state-Space) MARRS package and EM algorithm are discussed. The relative advantages and disadvantages of these approaches are highlighted. The performance of the vector autoregressive model based imputation methods is compared with that of three existing methods, namely, KNN, MARRRS and EM for dealing with Missing Data in multivariate time series, where an ECG Dataset is used as an a case study. The results show that VAR-IM produces a better recovering performance for Missing values than the other three methods. The advantages and limitations of the VAR-IM algorithm is also discussed.
-
Parametric and non-parametric methods to enhance prediction performance in the presence of Missing Data
2015 19th International Conference on System Theory Control and Computing (ICSTCC), 2015Co-Authors: Faraj Bashir, Hua-liang WeiAbstract:Most Missing Data Analysis techniques have focused on using model parameter estimation which depends on modern statistical Data Analysis methods such as maximum likelihood and multiple imputation. In fact, these modern methods are better than traditional methods (for example, complete Data Analysis and mean imputation approaches), and in many particular applications can give unbiased parametric estimation. Because these modern approaches depend on linear parametric regression, they do not give good results, especially if the Data distribution has highly nonlinear behaviour. This paper explains parametric estimation in cases of Missing Data, including an overview of parametric estimation with Missing Data, and provides accessible descriptions of nonlinear parametric and nonparametric estimation with Missing Data. In particular, this paper focuses on the effect of model selection methods on nonlinear parametric and nonparametric estimation in the presence of Missing Data. We also present Analysis of an example to illustrate the performance of the two methods.
-
MMAR - Model selection to enhance prediction performance in the presence of Missing Data
2015 20th International Conference on Methods and Models in Automation and Robotics (MMAR), 2015Co-Authors: Faraj Bashir, Hua-liang Wei, Abdollha M. BenomairAbstract:Most nonlinear modelling approaches focus on solving a model selection problem with complete Data, and in many cases original Data are pre-processed with some nonlinear transforms such as Box-Tidwell and fractional polynomial transformation. Often these approaches can lead to models that are better than traditional models (for example, logistic model and quadratic model). However, in the case of Missing Data, it is not easy to predict the relationship between the predictor and dependent variables; traditional nonlinear models in some cases of Missing Data Analysis give poor results. This paper explains nonlinear model selection techniques for Missing Data. It includes an overview of nonlinear model selection with complete Data, and provides accessible descriptions of Box-Tidwell and fractional polynomial methods for model selection. In particular, this paper focuses on a fractional polynomial method for nonlinear modelling in cases of Missing Data and presents Analysis examples to illustrate performance of the method.
-
ICPRAM (1) - Using Nonlinear Models to Enhance Prediction Performance with Incomplete Data
Proceedings of the International Conference on Pattern Recognition Applications and Methods, 2015Co-Authors: Faraj Bashir, Hua-liang WeiAbstract:A great deal of recent methodological research on Missing Data Analysis has focused on model parameter estimation using modern statistical methods such as maximum likelihood and multiple imputation. These approaches are better than traditional methods (for example listwise deletion and mean imputation methods). These modern techniques can lead to unbiased parametric estimation in many particular application cases. However, these methods do not work well in some cases especially for nonlinear systems that have highly nonlinear behaviour. This paper explains the linear parametric estimation in existence of Missing Data, which includes an overview of biased and unbiased linear parametric estimation with Missing Data, and provides accessible descriptions of expectation maximization (EM) algorithm and Gauss-Newton method. In particular, this paper proposes a Gauss-Newton iteration method for nonlinear parametric estimation in case of Missing Data. Since Gauss-Newton method needs initial values that are hard to obtain in the presence of Missing Data, the EM algorithm is thus used to estimate these initial values. In addition, we present two Analysis examples to illustrate the performance of the proposed methods.
Todd D. Little - One of the best experts on this subject based on the ideXlab platform.
-
Principled Missing Data Treatments
Prevention Science, 2018Co-Authors: Kyle M. Lang, Todd D. LittleAbstract:We review a number of issues regarding Missing Data treatments for intervention and prevention researchers. Many of the common Missing Data practices in prevention research are still, unfortunately, ill-advised (e.g., use of listwise and pairwise deletion, insufficient use of auxiliary variables). Our goal is to promote better practice in the handling of Missing Data. We review the current state of Missing Data methodology and recent Missing Data reporting in prevention research. We describe antiquated, ad hoc Missing Data treatments and discuss their limitations. We discuss two modern, principled Missing Data treatments: multiple imputation and full information maximum likelihood, and we offer practical tips on how to best employ these methods in prevention research. The principled Missing Data treatments that we discuss are couched in terms of how they improve causal and statistical inference in the prevention sciences. Our recommendations are firmly grounded in Missing Data theory and well-validated statistical principles for handling the Missing Data issues that are ubiquitous in biosocial and prevention research. We augment our broad survey of Missing Data Analysis with references to more exhaustive resources.
-
Handbook of Adolescent Psychology - Modeling Longitudinal Data from Research on Adolescence
Handbook of Adolescent Psychology, 2009Co-Authors: Todd D. Little, Noel A. Card, Kristopher J. Preacher, Elizabeth McconnellAbstract:Advantages of Longitudinal Data Design and Data Considerations Missing Data Analysis Techniques Panel Models Growth Curve Models Time Series and Related Models Mediation and Moderation in Longitudinal Data Conclusions Keywords: longitudinal Data; panel model; growth curve model; growth mixture model; time series model; mediation and moderation
Panagiota Chatzipetrou - One of the best experts on this subject based on the ideXlab platform.
-
A Framework of Statistical and Visualization Techniques for Missing Data Analysis in Software Cost Estimation
Intelligent Systems, 2018Co-Authors: Lefteris Angelis, Nikolaos Mittas, Panagiota ChatzipetrouAbstract:Software Cost Estimation (SCE) is a critical phase in software development projects. However, due to the growing complexity of the software itself, a common problem in building software cost models is that the available Datasets contain lots of Missing categorical Data. The purpose of this chapter is to show how a framework of statistical, computational, and visualization techniques can be used to evaluate and compare the effect of Missing Data techniques on the accuracy of cost estimation models. Hence, the authors use five Missing Data techniques: Multinomial Logistic Regression, Listwise Deletion, Mean Imputation, Expectation Maximization, and Regression Imputation. The evaluation and the comparisons are conducted using Regression Error Characteristic curves, which provide visual comparison of different prediction models, and Regression Error Operating Curves, which examine predictive power of models with respect to under- or over-estimation.
-
A Framework of Statistical and Visualization Techniques for Missing Data Analysis in Software Cost Estimation
Computer Systems and Software Engineering, 1Co-Authors: Lefteris Angelis, Nikolaos Mittas, Panagiota ChatzipetrouAbstract:Software Cost Estimation (SCE) is a critical phase in software development projects. However, due to the growing complexity of the software itself, a common problem in building software cost models ...
Yulei He - One of the best experts on this subject based on the ideXlab platform.
-
Missing Data Analysis using multiple imputation getting to the heart of the matter
Circulation-cardiovascular Quality and Outcomes, 2010Co-Authors: Yulei HeAbstract:Missing Data are a pervasive problem in health investigations. We describe some background of Missing Data Analysis and criticize ad hoc methods that are prone to serious problems. We then focus on multiple imputation, in which Missing cases are first filled in by several sets of plausible values to create multiple completed Datasets, then standard complete-Data procedures are applied to each completed Dataset, and finally the multiple sets of results are combined to yield a single inference. We introduce the basic concepts and general methodology and provide some guidance for application. For illustration, we use a study assessing the effect of cardiovascular diseases on hospice discussion for late stage lung cancer patients.