The Experts below are selected from a list of 88359 Experts worldwide ranked by ideXlab platform

Xihong Lin - One of the best experts on this subject based on the ideXlab platform.

  • Variance Component Testing in Generalized Linear Mixed Models for Longitudinal/Clustered Data and other Related Topics
    Random Effect and Latent Variable Model Selection, 2008
    Co-Authors: Daowen Zhang, Xihong Lin
    Abstract:

    Linear mixed models (Laird and Ware, 1982) and generalized linear mixed models (GLMMs) (Breslow and Clayton, 1993) have been widely used in many research areas, especially in the area of biomedical research, to analyze longitudinal and Clustered Data and multiple outcome Data. In a mixed effects model, subject-specific random effects are used to explicitly model between-subject variation in the Data and often assumed to follow a mean zero parametric distribution, e.g., multivariate normal, that depends on some unknown variance components. A large literature was developed in the last two decades for the estimation of regression coefficients and variance components in mixed effects models. See Diggle et al. (2002) and Verbeke and Molenberghs (2000, 2005) for an overview.

  • variance component testing in generalized linear mixed models for longitudinal Clustered Data and other related topics
    2008
    Co-Authors: Daowen Zhang, Xihong Lin
    Abstract:

    Linear mixed models (Laird and Ware, 1982) and generalized linear mixed models (GLMMs) (Breslow and Clayton, 1993) have been widely used in many research areas, especially in the area of biomedical research, to analyze longitudinal and Clustered Data and multiple outcome Data. In a mixed effects model, subject-specific random effects are used to explicitly model between-subject variation in the Data and often assumed to follow a mean zero parametric distribution, e.g., multivariate normal, that depends on some unknown variance components. A large literature was developed in the last two decades for the estimation of regression coefficients and variance components in mixed effects models. See Diggle et al. (2002) and Verbeke and Molenberghs (2000, 2005) for an overview.

  • efficient semiparametric marginal estimation for longitudinal Clustered Data
    Journal of the American Statistical Association, 2005
    Co-Authors: Naisyin Wang, Raymond J. Carroll, Xihong Lin
    Abstract:

    We consider marginal generalized semiparametric partially linear models for Clustered Data. Lin and Carroll derived the semiparametric efficient score function for this problem in the multivariate Gaussian case, but they were unable to construct a semiparametric efficient estimator that actually achieved the semiparametric information bound. Here we propose such an estimator and generalize the work to marginal generalized partially linear models. We investigate asymptotic relative efficiencies of the estimators that ignore the within-cluster correlation structure either in nonparametric curve estimation or throughout. We evaluate the finite-sample performance of these estimators through simulations and illustrate it using a longitudinal CD4 cell count Dataset. Both theoretical and numerical results indicate that properly taking into account the within-subject correlation among the responses can substantially improve efficiency.

  • Semiparametric regression for Clustered Data
    Biometrika, 2001
    Co-Authors: Xihong Lin, Raymond J. Carroll
    Abstract:

    We consider estimation in a semiparametric partially generalised linear model for Clustered Data using estimating equations. A marginal model is assumed where the mean of the outcome variable depends on some covariates parametrically and a cluster-level covariate nonparametrically. A profile-kernel method allowing for working correlation matrices is developed. We show that the nonparametric part of the model can be estimated using standard nonparametric methods, including smoothing-parameter estimation, and the parametric part of the model can be estimated in a profile fashion. The asymptotic distributions of the parameter estimators are derived, and the optimal estimators of both the nonparametric and parametric parts are shown to be obtained when the working correlation matrix equals the actual correlation matrix. The asymptotic covariance matrix of the parameter estimator is consistently estimated by the sandwich estimator. We show that the semiparametric efficient score takes on a simple form and our profile-kernel method is semiparametric efficient. The results for the case where the nonparametric part of the model is an observation-level covariate are noted to be dramatically different.

  • Semiparametric Regression for Clustered Data Using Generalized Estimating Equations
    Journal of the American Statistical Association, 2001
    Co-Authors: Xihong Lin, Raymond J. Carroll
    Abstract:

    We consider estimation in a semiparametric generalized linear model for Clustered Data using estimating equations. Our results apply to the case where the number of observations per cluster is finite, whereas the number of clusters is large. The mean of the outcome variable μ is of the form g(μ) = XTβ + θ(T), where g(·) is a link function, X and T are covariates, β is an unknown parameter vector, and θ(t) is an unknown smooth function. Kernel estimating equations proposed previously in the literature are used to estimate the infinite-dimensional nonparametric function θ(t), and a profile-based estimating equation is used to estimate the finite-dimensional parameter vector β. We show that for Clustered Data, this conventional profile-kernel method often fails to yield a √n-consistent estimator of β along with appropriate inference unless working independence is assumed or θ(t) is artificially undersmoothed, in which case asymptotic inference is possible. To gain insight into these results, we derive the se...

Bernard Rosner - One of the best experts on this subject based on the ideXlab platform.

  • Wilcoxon Rank-Based Tests for Clustered Data with R Package clusrank
    Journal of Statistical Software, 2020
    Co-Authors: Yujing Jiang, Bernard Rosner, Mei-ling Ting Lee, Jun Yan
    Abstract:

    Wilcoxon rank-based tests are distribution-free alternatives to the popular two-sample and paired t tests. For independent Data, they are available in several R packages such as stats and coin. For Clustered Data, in spite of the recent methodological developments, there did not exist an R package that makes them available at one place. We present a package clusrank where the latest developments are implemented and wrapped under a unified user-friendly interface. With different methods dispatched based on the inputs, this package offers great flexibility in rank-based tests for various Clustered Data. Exact tests based on permutations are also provided for some methods. Details of the major schools of different methods are briefly reviewed. Usages of the package clusrank are illustrated with simulated Data as well as a real Dataset from an ophthalmological study. The package also enables convenient comparison between selected methods under settings that have not been studied before and the results are discussed.

  • Estimation of rank correlation for Clustered Data.
    Statistics in medicine, 2017
    Co-Authors: Bernard Rosner, Robert J. Glynn
    Abstract:

    It is well known that the sample correlation coefficient (Rxy ) is the maximum likelihood estimator of the Pearson correlation (ρxy ) for independent and identically distributed (i.i.d.) bivariate normal Data. However, this is not true for ophthalmologic Data where X (e.g., visual acuity) and Y (e.g., visual field) are available for each eye and there is positive intraclass correlation for both X and Y in fellow eyes. In this paper, we provide a regression-based approach for obtaining the maximum likelihood estimator of ρxy for Clustered Data, which can be implemented using standard mixed effects model software. This method is also extended to allow for estimation of partial correlation by controlling both X and Y for a vector U_ of other covariates. In addition, these methods can be extended to allow for estimation of rank correlation for Clustered Data by (i) converting ranks of both X and Y to the probit scale, (ii) estimating the Pearson correlation between probit scores for X and Y, and (iii) using the relationship between Pearson and rank correlation for bivariate normally distributed Data. The validity of the methods in finite-sized samples is supported by simulation studies. Finally, two examples from ophthalmology and analgesic abuse are used to illustrate the methods. Copyright © 2017 John Wiley & Sons, Ltd.

  • The Wilcoxon Signed Rank Test for Paired Comparisons of Clustered Data
    Biometrics, 2005
    Co-Authors: Bernard Rosner, Robert J. Glynn, Mei-ling Ting Lee
    Abstract:

    The Wilcoxon signed rank test is a frequently used nonparametric test for paired Data (e.g., consisting of pre- and posttreatment measurements) based on independent units of analysis. This test cannot be used for paired comparisons arising from Clustered Data (e.g., if paired comparisons are available for each of two eyes of an individual). To incorporate clustering, a generalization of the randomization test formulation for the signed rank test is proposed, where the unit of randomization is at the cluster level (e.g., person), while the individual paired units of analysis are at the subunit within cluster level (e.g., eye within person). An adjusted variance estimate of the signed rank test statistic is then derived, which can be used for either balanced (same number of subunits per cluster) or unbalanced (different number of subunits per cluster) Data, with an exchangeable correlation structure, with or without tied values. The resulting test statistic is shown to be asymptotically normal as the number of clusters becomes large, if the cluster size is bounded. Simulation studies are performed based on simulating correlated ranked Data from a signed log-normal distribution. These studies indicate appropriate type I error for Data sets with > or =20 clusters and a superior power profile compared with either the ordinary signed rank test based on the average cluster difference score or the multivariate signed rank test of Puri and Sen. Finally, the methods are illustrated with two Data sets, (i) an ophthalmologic Data set involving a comparison of electroretinogram (ERG) Data in retinitis pigmentosa (RP) patients before and after undergoing an experimental surgical procedure, and (ii) a nutritional Data set based on a randomized prospective study of nutritional supplements in RP patients where vitamin E intake outside of study capsules is compared before and after randomization to monitor compliance with nutritional protocols.

  • use of the mann whitney u test for Clustered Data
    Statistics in Medicine, 1999
    Co-Authors: Bernard Rosner, D Grove
    Abstract:

    The Mann-Whitney U-test is ubiquitous in statistical practice for the comparison of measures of location for two samples where the assumption of normality is questionable. Frequently, one has replicate Data for each individual in a group and would like to compare measures of central tendency between groups without assuming normality. For this purpose, we present a generalization of the Mann-Whitney U-test for Clustered Data. The test is performed by computing zc = (Wc - mu c)/sigma c, approximately N(0, 1) under H0, where Wc, mu c are the observed and expected Mann-Whitney U-statistic based on a comparison of all pairs of replicates in the two groups and sigma c is the standard deviation of Wc that is modified to account for clustering effects within a cluster. We obtain an explicit variance formula that is a function of four clustering parameters. We validate the properties of the test procedure in a simulation study. We illustrate the methods with an example comparing the baseline Humphrey visual field between two treatment groups in a randomized clinical trial of patients with retinitis pigmentosa (RP).

Jeffrey S. Simonoff - One of the best experts on this subject based on the ideXlab platform.

  • Unbiased regression trees for longitudinal and Clustered Data
    Computational Statistics & Data Analysis, 2015
    Co-Authors: Jeffrey S. Simonoff
    Abstract:

    A new version of the RE-EM regression tree method for longitudinal and Clustered Data is presented. The RE-EM tree is a methodology that combines the structure of mixed effects models for longitudinal and Clustered Data with the flexibility of tree-based estimation methods. The RE-EM tree is less sensitive to parametric assumptions and provides improved predictive power compared to linear models with random effects and regression trees without random effects. The previously-suggested methodology used the CART tree algorithm for tree building, and therefore that RE-EM regression tree method inherits the tendency of CART to split on variables with more possible split points at the expense of those with fewer split points. A revised version of the RE-EM regression tree corrects for this bias by using the conditional inference tree as the underlying tree algorithm instead of CART. Simulation studies show that the new version is indeed unbiased, and has several improvements over the original RE-EM regression tree in terms of prediction accuracy and the ability to recover the correct tree structure.

  • Unbiased Regression Trees for Longitudinal and Clustered Data
    SSRN Electronic Journal, 2014
    Co-Authors: Jeffrey S. Simonoff
    Abstract:

    This paper presents a new version of the RE-EM regression tree method for longitudinal and Clustered Data. The RE-EM tree is a methodology that combines the structure of mixed effects models for longitudinal and Clustered Data with the flexibility of tree-based estimation methods. The RE-EM tree is less sensitive to parametric assumptions and provides improved predictive power compared to linear models with random effects and regression trees without random effects. The previously-suggested methodology used the CART tree algorithm for tree building, and therefore that RE-EM regression tree method inherits the tendency of CART to split on variables with more possible split points at the expense of those with fewer split points. A revised version of the RE-EM regression tree corrects for this bias by using the conditional inference tree as the underlying tree algorithm instead of CART. Simulation studies show that the new version is indeed unbiased, and has several improvements over the original RE-EM regression tree in terms of prediction accuracy and the ability to recover the correct true structure.

  • RE-EM trees: a Data mining approach for longitudinal and Clustered Data
    Machine Learning, 2012
    Co-Authors: Rebecca J. Sela, Jeffrey S. Simonoff
    Abstract:

    Longitudinal Data refer to the situation where repeated observations are available for each sampled object. Clustered Data, where observations are nested in a hierarchical structure within objects (without time necessarily being involved) represent a similar type of situation. Methodologies that take this structure into account allow for the possibilities of systematic differences between objects that are not related to attributes and autocorrelation within objects across time periods. A standard methodology in the statistics literature for this type of Data is the mixed effects model, where these differences between objects are represented by so-called “random effects” that are estimated from the Data (population-level relationships are termed “fixed effects,” together resulting in a mixed effects model). This paper presents a methodology that combines the structure of mixed effects models for longitudinal and Clustered Data with the flexibility of tree-based estimation methods. We apply the resulting estimation method, called the RE-EM tree, to pricing in online transactions, showing that the RE-EM tree is less sensitive to parametric assumptions and provides improved predictive power compared to linear models with random effects and regression trees without random effects. We also apply it to a smaller Data set examining accident fatalities, and show that the RE-EM tree strongly outperforms a tree without random effects while performing comparably to a linear model with random effects. We also perform extensive simulation experiments to show that the estimator improves predictive performance relative to regression trees without random effects and is comparable or superior to using linear models with random effects in more general situations.

Somnath Datta - One of the best experts on this subject based on the ideXlab platform.

  • Inferring marginal association with paired and unpaired Clustered Data.
    Statistical methods in medical research, 2016
    Co-Authors: Douglas J. Lorenz, Steven M. Levy, Somnath Datta
    Abstract:

    In the marginal analysis of Clustered Data, where the marginal distribution of interest is that of a typical observation within a typical cluster, analysis by reweighting has been introduced as a useful tool for estimating parameters of these marginal distributions. Such reweighting methods have foundation in within-cluster resampling schemes that marginalize potential informativeness due to cluster size or within-cluster covariate distribution, to which reweighting methods are asymptotically equivalent. In this paper, we introduce a reweighting scheme for the marginal analysis of Clustered Data that generalizes prior reweighting methods, with a particular application to measuring bivariate correlation in unpaired Clustered Data, in which observations of two random variables are not naturally paired at the within-cluster level. We develop unpaired Clustered Data analogs of well-known product moment correlation coefficients (Pearson, Spearman, phi), as well as the polyserial coefficient for measuring correlation between one discrete and one continuous variable. We evaluate the performance of these coefficients via a simulation study and demonstrate their use by finding no statistically significant association between dental caries at an early age and dental fluorosis at age 13 using a large dental Dataset.

  • a rank sum test for Clustered Data when the number of subjects in a group within a cluster is informative
    Biometrics, 2016
    Co-Authors: Sandipan Dutta, Somnath Datta
    Abstract:

    The Wilcoxon rank-sum test is a popular nonparametric test for comparing two independent populations (groups). In recent years, there have been renewed attempts in extending the Wilcoxon rank sum test for Clustered Data, one of which (Datta and Satten, 2005, Journal of the American Statistical Association 100, 908-915) addresses the issue of informative cluster size, i.e., when the outcomes and the cluster size are correlated. We are faced with a situation where the group specific marginal distribution in a cluster depends on the number of observations in that group (i.e., the intra-cluster group size). We develop a novel extension of the rank-sum test for handling this situation. We compare the performance of our test with the Datta-Satten test, as well as the naive Wilcoxon rank sum test. Using a naturally occurring simulation model of informative intra-cluster group size, we show that only our test maintains the correct size. We also compare our test with a classical signed rank test based on averages of the outcome values in each group paired by the cluster membership. While this test maintains the size, it has lower power than our test. Extensions to multiple group comparisons and the case of clusters not having samples from all groups are also discussed. We apply our test to determine whether there are differences in the attachment loss between the upper and lower teeth and between mesial and buccal sites of periodontal patients.

  • Robust estimation of marginal regression parameters in Clustered Data.
    Statistical modelling, 2014
    Co-Authors: Somnath Datta, James D. Beck
    Abstract:

    We develop robust methods for analyzing Clustered Data where estimation of marginal regression parameters is of interest. Inverse cluster size reweighting in the objective function to be minimized is incorporated to handle the issue of informative cluster size. Performance of the resulting estimators is studied by simulation. Large sample inference and variance estimation is carried out. The methodology is illustrated using a periodontal disease Dataset.

  • A General Class of Signed Rank Tests for Clustered Data when the Cluster Size is Potentially Informative.
    Journal of nonparametric statistics, 2012
    Co-Authors: Somnath Datta, Jaakko Nevalainen, Hannu Oja
    Abstract:

    Rank-based tests are alternatives to likelihood-based tests popularised by their relative robustness and underlying elegant mathematical theory. There has been a surge in research activities in this area in recent years since a number of researchers are working to develop and extend rank-based procedures to Clustered-dependent Data which include situations with known correlation structures (e.g. as in mixed effects models) as well as more general form of dependence. The purpose of this paper is to test the symmetry of a marginal distribution under Clustered Data. However, unlike most other papers in the area, we consider the possibility that the cluster size is a random variable whose distribution is dependent on the distribution of the variable of interest within a cluster. This situation typically arises when the clusters are defined in a natural way (e.g. not controlled by the experimenter or statistician) and in which the size of the cluster may carry information about the distribution of Data values ...

  • Marginal association measures for Clustered Data
    Statistics in medicine, 2011
    Co-Authors: Douglas J. Lorenz, Somnath Datta, Susan J. Harkema
    Abstract:

    The use of correlation coefficients in measuring the association between two continuous variables is common, but regular methods of calculating correlations have not been extended to the Clustered Data framework. For Clustered Data in which observations within a cluster may be correlated, regular inferential procedures for calculating marginal association between two variables can be biased. This is particularly true for Data in which the number of observations in a given cluster is informative for the association being measured. In this paper, we apply the principle of inverse cluster size reweighting to develop estimators of marginal correlation that remain valid in the Clustered Data framework when cluster size is informative for the correlation being measured. These correlations are derived as analogs to regular correlation estimators for continuous, independent Data, namely, Pearson’s ρ and Kendall’s τ. We present the results of a simple simulation study demonstrating the appropriateness of our proposed estimators and the inherent bias of other inferential procedures for Clustered Data. We illustrate their use through an application to Data from patients with incomplete spinal cord injury in the USA.

Yvonne Vergouwe - One of the best experts on this subject based on the ideXlab platform.

  • a simulation study of sample size demonstrated the importance of the number of events per variable to develop prediction models in Clustered Data
    Journal of Clinical Epidemiology, 2015
    Co-Authors: L Wynants, Walter Bouwmeester, K G M Moons, Mirjam Moerbeek, Dirk Timmerman, S Van Huffel, B Van Calster, Yvonne Vergouwe
    Abstract:

    Abstract Objectives This study aims to investigate the influence of the amount of clustering [intraclass correlation (ICC) = 0%, 5%, or 20%], the number of events per variable (EPV) or candidate predictor (EPV = 5, 10, 20, or 50), and backward variable selection on the performance of prediction models. Study Design and Setting Researchers frequently combine Data from several centers to develop clinical prediction models. In our simulation study, we developed models from Clustered training Data using multilevel logistic regression and validated them in external Data. Results The amount of clustering was not meaningfully associated with the models' predictive performance. The median calibration slope of models built in samples with EPV = 5 and strong clustering (ICC = 20%) was 0.71. With EPV = 5 and ICC = 0%, it was 0.72. A higher EPV related to an increased performance: the calibration slope was 0.85 at EPV = 10 and ICC = 20% and 0.96 at EPV = 50 and ICC = 20%. Variable selection sometimes led to a substantial relative bias in the estimated predictor effects (up to 118% at EPV = 5), but this had little influence on the model's performance in our simulations. Conclusion We recommend at least 10 EPV to fit prediction models in Clustered Data using logistic regression. Up to 50 EPV may be needed when variable selection is performed.

  • assessing discriminative ability of risk models in Clustered Data
    BMC Medical Research Methodology, 2014
    Co-Authors: David Van Klaveren, Ewout W Steyerberg, Pablo Perel, Yvonne Vergouwe
    Abstract:

    The discriminative ability of a risk model is often measured by Harrell’s concordance-index (c-index). The c-index estimates for two randomly chosen subjects the probability that the model predicts a higher risk for the subject with poorer outcome (concordance probability). When Data are Clustered, as in multicenter Data, two types of concordance are distinguished: concordance in subjects from the same cluster (within-cluster concordance probability) and concordance in subjects from different clusters (between-cluster concordance probability). We argue that the within-cluster concordance probability is most relevant when a risk model supports decisions within clusters (e.g. who should be treated in a particular center). We aimed to explore different approaches to estimate the within-cluster concordance probability in Clustered Data. We used Data of the CRASH trial (2,081 patients Clustered in 35 centers) to develop a risk model for mortality after traumatic brain injury. To assess the discriminative ability of the risk model within centers we first calculated cluster-specific c-indexes. We then pooled the cluster-specific c-indexes into a summary estimate with different meta-analytical techniques. We considered fixed effect meta-analysis with different weights (equal; inverse variance; number of subjects, events or pairs) and random effects meta-analysis. We reflected on pooling the estimates on the log-odds scale rather than the probability scale. The cluster-specific c-index varied substantially across centers (IQR = 0.70-0.81; I 2 = 0.76 with 95% confidence interval 0.66 to 0.82). Summary estimates resulting from fixed effect meta-analysis ranged from 0.75 (equal weights) to 0.84 (inverse variance weights). With random effects meta-analysis – accounting for the observed heterogeneity in c-indexes across clusters – we estimated a mean of 0.77, a between-cluster variance of 0.0072 and a 95% prediction interval of 0.60 to 0.95. The normality assumptions for derivation of a prediction interval were better met on the probability than on the log-odds scale. When assessing the discriminative ability of risk models used to support decisions at cluster level we recommend meta-analysis of cluster-specific c-indexes. Particularly, random effects meta-analysis should be considered.

  • prediction models for Clustered Data comparison of a random intercept and standard regression model
    BMC Medical Research Methodology, 2013
    Co-Authors: Walter Bouwmeester, Jos W R Twisk, Teus H Kappen, Wilton A Van Klei, Karel G M Moons, Yvonne Vergouwe
    Abstract:

    Background: When study Data are Clustered, standard regression analysis is considered inappropriate and analytical techniques for Clustered Data need to be used. For prediction research in which the interest of predictor effects is on the patient level, random effect regression models are probably preferred over standard regression analysis. It is well known that the random effect parameter estimates and the standard logistic regression parameter estimates are different. Here, we compared random effect and standard logistic regression models for their ability to provide accurate predictions. Methods: Using an empirical study on 1642 surgical patients at risk of postoperative nausea and vomiting, who were treated by one of 19 anesthesiologists (clusters), we developed prognostic models either with standard or random intercept logistic regression. External validity of these models was assessed in new patients from other anesthesiologists. We supported our results with simulation studies using intra-class correlation coefficients (ICC) of 5%, 15%, or 30%. Standard performance measures and measures adapted for the Clustered Data structure were estimated. Results: The model developed with random effect analysis showed better discrimination than the standard approach, if the cluster effects were used for risk prediction (standard c-index of 0.69 versus 0.66). In the external validation set, both models showed similar discrimination (standard c-index 0.68 versus 0.67). The simulation study confirmed these results. For Datasets with a high ICC (≥15%), model calibration was only adequate in external subjects, if the used performance measure assumed the same Data structure as the model development method: standard calibration measures showed good calibration for the standard developed model, calibration measures adapting the Clustered Data structure showed good calibration for the prediction model with random intercept. Conclusion: The models with random intercept discriminate better than the standard model only if the cluster effect is used for predictions. The prediction model with random intercept had good calibration within clusters.