The Experts below are selected from a list of 48 Experts worldwide ranked by ideXlab platform
Andreas Hense - One of the best experts on this subject based on the ideXlab platform.
-
statistical downscaling of extreme precipitation events using censored quantile regression
Monthly Weather Review, 2007Co-Authors: Petra Friederichs, Andreas HenseAbstract:Abstract A statistical downscaling approach for extremes using censored quantile regression is presented. Conditional quantiles of station data (e.g., daily precipitation sums) in Germany are estimated by means of the large-scale circulation as represented by the NCEP reanalysis data. It is shown that a mixed discrete–Continuous Response Variable, such as a daily precipitation sum, can be statistically modeled by a censored Variable. Furthermore, a conditional quantile skill score is formulated to assess the relative gain of a quantile forecast compared with a reference forecast. Just like multiple regression for expectation values, quantile regression provides a tool to formulate a model output statistics system for extremal quantiles.
Yue Shentu - One of the best experts on this subject based on the ideXlab platform.
-
A note on dichotomization of Continuous Response Variable in the presence of contamination and model misspecification
Statistics in Medicine, 2010Co-Authors: Yue ShentuAbstract:The purpose of this note is to raise awareness of the complexity of the practice involving dichotomization. It is well known that the regular regression models are effective tools for analyzing Gaussian-type Response Variables, and researchers are often told that it is a 'bad idea' to practice dichotomization if Continuous measurements are available. We demonstrate through special cases, however, that there is another side of the story if the Response Variable is contaminated. Although dichotomization causes loss of information, it can also reduce input of contamination. If the reduction of contamination input outweighs the loss of information, analysis based on dichotomization can sometimes provide better results. We derive formulas of bias and variance for binary regression estimators under a contamination model of unknown additive errors, and compare them with both the least squares and robust M-estimators from the corresponding linear regression analysis using Continuous Responses. As a case study, we study extensively the case in which the observed Response is contaminated by an error with a mean and a variance proportional to the mean and the variance of the uncontaminated true Response. Conditions under which dichotomization is preferred are obtained. A simulation study based on a real data setting is provided, which supports the theoretical developments.
Petra Friederichs - One of the best experts on this subject based on the ideXlab platform.
-
statistical downscaling of extreme precipitation events using censored quantile regression
Monthly Weather Review, 2007Co-Authors: Petra Friederichs, Andreas HenseAbstract:Abstract A statistical downscaling approach for extremes using censored quantile regression is presented. Conditional quantiles of station data (e.g., daily precipitation sums) in Germany are estimated by means of the large-scale circulation as represented by the NCEP reanalysis data. It is shown that a mixed discrete–Continuous Response Variable, such as a daily precipitation sum, can be statistically modeled by a censored Variable. Furthermore, a conditional quantile skill score is formulated to assess the relative gain of a quantile forecast compared with a reference forecast. Just like multiple regression for expectation values, quantile regression provides a tool to formulate a model output statistics system for extremal quantiles.
Francois Caron - One of the best experts on this subject based on the ideXlab platform.
-
scalable bayesian nonparametric regression via a plackett luce model for conditional ranks
Electronic Journal of Statistics, 2016Co-Authors: Tristan Graydavies, Christopher Holmes, Francois CaronAbstract:We present a novel Bayesian nonparametric regression model for covariates X and Continuous Response Variable Y ∈ ℝ. The model is parametrized in terms of marginal distributions for Y and X and a regression function which tunes the stochastic ordering of the conditional distributions F (y|x). By adopting an approximate composite likelihood approach, we show that the resulting posterior inference can be decoupled for the separate components of the model. This procedure can scale to very large datasets and allows for the use of standard, existing, software from Bayesian nonparametric density estimation and Plackett-Luce ranking estimation to be applied. As an illustration, we show an application of our approach to a US Census dataset, with over 1,300,000 data points and more than 100 covariates.
Jingshiang Hwang - One of the best experts on this subject based on the ideXlab platform.
-
stepwise paring down variation for identifying influential multi factor interactions related to a Continuous Response Variable
Statistics in Biosciences, 2012Co-Authors: Jingshiang HwangAbstract:Although several model-based methods are promising for the identification of influential single factors and multi-factor interactions, few are widely used in real applications for most of the model-selection procedures are complex and/or infeasible in computation for high-dimensional data. In particular, the ability of the methods to reveal more true factors and fewer false ones often relies heavily on the selection of appropriate values of tuning parameters, which is still a difficult task to practical analysts. This article provides a simple algorithm modified from stepwise forward regression for the identification of influential factors. Instead of keeping the identified factors in the next models for adjustment in stepwise regression, we propose to subtract the effects of identified factors in each run and always fit a single-term model to the effect-subtracted Responses. The computation is lighter as the proposed method only involves calculations of a simple test statistic; and therefore it could be applied to screen ultrahigh-dimensional data for important single factors and multi-factor interactions. Most importantly, we have proposed a novel stopping rule of using a constant threshold for the simple test statistic, which is different from the conventional stepwise regression with AIC or BIC criterion. The performance of the new algorithm has been confirmed competitive by extensive simulation studies compared to several methods available in R packages, including the popular group lasso, surely independence screening, Bayesian quantitative trait locus mapping methods and others. Findings from two real data examples, including a genome-wide association study, demonstrate additional useful information of high-order interactions that can be gained from implementing the proposed algorithm.