The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform
Jaakko Makinen - One of the best experts on this subject based on the ideXlab platform.
-
a bound for the euclidean norm of the difference between the best linear unbiased estimator and a linear unbiased estimator
Journal of Geodesy, 2002Co-Authors: Jaakko MakinenAbstract:A bound is established for the Euclidean norm of the difference between the best linear unbiased estimator and any linear unbiased estimator in the general linear model. The bound involves the spectral norm of the difference between the dispersion matrices of the two estimators, and the Residual Sum of squares, all evaluated at the asSumed model, but is independent of the provenance of the observation vector at hand. The bound, a straightforward consequence of first principles in Gauss–Markov theory, generalizes previous results on the difference between the best linear unbiased estimator and the ordinary least-squares estimator. In a numerical example from repeated precise levelling, the bound is used to analyse the sensitivity of estimates of vertical motion to the choice of estimator.
Stanley Lemeshow - One of the best experts on this subject based on the ideXlab platform.
-
a comparison of goodness of fit tests for the logistic regression model
Statistics in Medicine, 1997Co-Authors: David W Hosmer, Trina Hosmer, Le S Cessie, Stanley LemeshowAbstract:Recent work has shown that there may be disadvantages in the use of the chi-square-like goodness-of-fit tests for the logistic regression model proposed by Hosmer and Lemeshow that use fixed groups of the estimated probabilities. A particular concern with these grouping strategies based on estimated probabilities, fitted values, is that groups may contain subjects with widely different values of the covariates. It is possible to demonstrate situations where one set of fixed groups shows the model fits while the test rejects fit using a different set of fixed groups. We compare the performance by simulation of these tests to tests based on smoothed Residuals proposed by le Cessie and Van Houwelingen and Royston, a score test for an extended logistic regression model proposed by Stukel, the Pearson chi-square and the unweighted Residual Sum-of-squares. These simulations demonstrate that all but one of Royston's tests have the correct size. An examination of the performance of the tests when the correct model has a quadratic term but a model containing only the linear term has been fit shows that the Pearson chi-square, the unweighted Sum-of-squares, the Hosmer-Lemeshow decile of risk, the smoothed Residual Sum-of-squares and Stukel's score test, have power exceeding 50 per cent to detect moderate departures from linearity when the sample size is 100 and have power over 90 per cent for these same alternatives for samples of size 500. All tests had no power when the correct model had an interaction between a dichotomous and continuous covariate but only the continuous covariate model was fit. Power to detect an incorrectly specified link was poor for samples of size 100. For samples of size 500 Stukel's score test had the best power but it only exceeded 50 per cent to detect an asymmetric link function. The power of the unweighted Sum-of-squares test to detect an incorrectly specified link function was slightly less than Stukel's score test. We illustrate the tests within the context of a model for factors associated with low birth weight.
-
a comparison of goodness of fit tests for the logistic regression model
Statistics in Medicine, 1997Co-Authors: David W Hosmer, Trina Hosmer, Le S Cessie, Stanley LemeshowAbstract:SumMARY Recent work has shown that there may be disadvantages in the use of the chi-square-like goodness-of-fit tests for the logistic regression model proposed by Hosmer and Lemeshow that use fixed groups of the estimated probabilities. A particular concern with these grouping strategies based on estimated probabilities, fitted values, is that groups may contain subjects with widely di⁄erent values of the covariates. It is possible to demonstrate situations where one set of fixed groups shows the model fits while the test rejects fit using a di⁄erent set of fixed groups. We compare the performance by simulation of these tests to tests based on smoothed Residuals proposed by le Cessie and Van Houwelingen and Royston, a score test for an extended logistic regression model proposed by Stukel, the Pearson chi-square and the unweighted Residual Sum-of- squares. These simulations demonstrate that all but one of Royston’s tests have the correct size. An examination of the performance of the tests when the correct model has a quadratic term but a model containing only the linear term has been fit shows that the Pearson chi-square, the unweighted Sum-ofsquares, the Hosmer—Lemeshow decile of risk, the smoothed Residual Sum-of-squares and Stukel’s score test, have power exceeding 50 per cent to detect moderate departures from linearity when the sample size is 100 and have power over 90 per cent for these same alternatives for samples of size 500. All tests had no power when the correct model had an interaction between a dichotomous and continuous covariate but only the continuous covariate model was fit. Power to detect an incorrectly specified link was poor for samples of size 100. For samples of size 500 Stukel’s score test had the best power but it only exceeded 50 per cent to detect an asymmetric link function. The power of the unweighted Sum-of-squares test to detect an incorrectly specified link function was slightly less than Stukel’s score test. We illustrate the tests within the context of a model for factors associated with low birth weight. ( 1997 by John Wiley & Sons, Ltd. Stat. Med., Vol. 16, 965—980 (1997).
Kirthevasan Kandasamy - One of the best experts on this subject based on the ideXlab platform.
-
additive approximations in high dimensional nonparametric regression via the salsa
International Conference on Machine Learning, 2016Co-Authors: Kirthevasan KandasamyAbstract:High dimensional nonparametric regression is an inherently difficult problem with known lower bounds depending exponentially in dimension. A popular strategy to alleviate this curse of dimensionality has been to use additive models of first order, which model the regression function as a Sum of independent functions on each dimension. Though useful in controlling the variance of the estimate, such models are often too restrictive in practical settings. Between non-additive models which often have large variance and first order additive models which have large bias, there has been little work to exploit the trade-off in the middle via additive models of intermediate order. In this work, we propose SALSA, which bridges this gap by allowing interactions between variables, but controls model capacity by limiting the order of interactions. SALSA minimises the Residual Sum of squares with squared RKHS norm penalties. Algorithmically, it can be viewed as Kernel Ridge Regression with an additive kernel. When the regression function is additive, the excess risk is only polynomial in dimension. Using the Girard-Newton formulae, we efficiently Sum over a combinatorial number of terms in the additive expansion. Via a comparison on 15 real datasets, we show that our method is competitive against 21 other alternatives.
-
additive approximations in high dimensional nonparametric regression via the salsa
arXiv: Machine Learning, 2016Co-Authors: Kirthevasan KandasamyAbstract:High dimensional nonparametric regression is an inherently difficult problem with known lower bounds depending exponentially in dimension. A popular strategy to alleviate this curse of dimensionality has been to use additive models of \emph{first order}, which model the regression function as a Sum of independent functions on each dimension. Though useful in controlling the variance of the estimate, such models are often too restrictive in practical settings. Between non-additive models which often have large variance and first order additive models which have large bias, there has been little work to exploit the trade-off in the middle via additive models of intermediate order. In this work, we propose SALSA, which bridges this gap by allowing interactions between variables, but controls model capacity by limiting the order of interactions. SALSA minimises the Residual Sum of squares with squared RKHS norm penalties. Algorithmically, it can be viewed as Kernel Ridge Regression with an additive kernel. When the regression function is additive, the excess risk is only polynomial in dimension. Using the Girard-Newton formulae, we efficiently Sum over a combinatorial number of terms in the additive expansion. Via a comparison on $15$ real datasets, we show that our method is competitive against $21$ other alternatives.
Shahjahan Khan - One of the best experts on this subject based on the ideXlab platform.
-
Optimal tolerance regions for future regression vector and Residual Sum of squares of multiple regression model with multivariate spherically contoured errors
Statistical Papers, 2007Co-Authors: Shahjahan KhanAbstract:This paper considers multiple regression model with multivariate spherically symmetric errors to determine optimal β-expectation tolerance regions for the future regression vector (FRV) and future Residual Sum of squares (FRSS) by using the prediction distributions of some appropriate functions of future responses. The prediction distribution of the FRV, conditional on the observed responses, is multivariate Student-t distribution. Similarly, the prediction distribution of the FRSS is a beta distribution. The optimal β-expectation tolerance regions for the FRV and FRSS have been obtained based on the F -distribution and beta distribution, respectively. The results in this paper are applicable for multiple regression model with normal and Student-t errors.
-
prediction distribution of future regression and Residual Sum of squares matrices for multivariate simple regression model with correlated normal responses
2006Co-Authors: Shahjahan KhanAbstract:This paper considers multivariate simple regression model under normally distributed errors, for both realized and future responses, with unknown regression parameters and covariance matrix. The prediction distributions of the future regression matrix (FRM) and future Residual Sum of squares matrix (FRSSM) for the future regression model are obtained. Conditional on the realized responses, the FRM follows a matrix T distribution whose shape parameter depends on the sample size and the dimension of the regression parameters in the model, and the FRSSM follows a scaled generalized beta distribution. The same results have been obtained by both the classical and Bayesian methods under uniform prior.
-
distribution of future location vector and Residual Sum of squares for multivariate location scale model with spherically contoured errors
2006Co-Authors: Shahjahan Khan, Enamul KabirAbstract:The multivariate location-scale model with a family of spherically contoured errors is considered for both realized and future responses. The predictive distributions of the future location vector (FLV) and future Residual Sum of squares (FRSS) for the future responses are obtained. Conditional on the realized responses, the FLV follows a multivariate Student-t distribution whose shape parameter depends on the sample size and the dimension of the location parameters of the model, and the FRSS follows a scaled beta distribution. The results obtained by both the classical and Bayesian methods under uniform prior are identical. This paper generalizes the results for location-scale models with multivariate normal and Student-t models to a wider family of spherically/ellipticcally contoured models.
-
predictive distribution of regression vector and Residual Sum of squares for normal multiple regression model
Communications in Statistics-theory and Methods, 2005Co-Authors: Shahjahan KhanAbstract:This article proposes predictive inference for the multiple regression model with independent normal errors. The distributions of the sample regression vector (SRV) and the Residual Sum of squares (RSS) for the model are derived by using invariant differentials. Also, the predictive distributions of the future regression vector (FRV) and the future Residual Sum of squares (FRSS) for the future regression model are obtained. Conditional on the realized responses, the FRV is found to follow a multivariate Student t distribution, and that of the Residual Sum of squares follows a scaled beta distribution. The new results have been applied to the market return and accounting rate data to illustrate its application.
Robert Tibshirani - One of the best experts on this subject based on the ideXlab platform.
-
a study of error variance estimation in lasso regression
Statistica Sinica, 2016Co-Authors: Stephen Reid, Robert Tibshirani, J FriedmanAbstract:Variance estimation in the linear model when p > n is a difficult problem. Standard least squares estimation techniques do not apply. Several variance estimators have been proposed in the literature, all with accompanying asymptotic results proving consistency and asymptotic normality under a variety of asSumptions. It is found, however, that most of these estimators suffer large biases in finite samples when true underlying signals become less sparse with larger per element signal strength. One estimator seems to merit more attention than it has received in the literature: a Residual Sum of squares based estimator using Lasso coefficients with regularisation parameter selected adaptively (via cross-validation). In this paper, we review several variance estimators and perform a reasonably extensive simulation study in an attempt to compare their finite sample performance. It would seem from the results that variance estimators with adaptively chosen regularisation parameters perform admirably over a broad range of sparsity and signal strength settings. Finally, some intial theoretical analyses pertaining to these types of estimators are proposed and developed.
-
a study of error variance estimation in lasso regression
arXiv: Methodology, 2013Co-Authors: Stephen Reid, Robert Tibshirani, J FriedmanAbstract:Variance estimation in the linear model when $p > n$ is a difficult problem. Standard least squares estimation techniques do not apply. Several variance estimators have been proposed in the literature, all with accompanying asymptotic results proving consistency and asymptotic normality under a variety of asSumptions. It is found, however, that most of these estimators suffer large biases in finite samples when true underlying signals become less sparse with larger per element signal strength. One estimator seems to be largely neglected in the literature: a Residual Sum of squares based estimator using Lasso coefficients with regularisation parameter selected adaptively (via cross-validation). In this paper, we review several variance estimators and perform a reasonably extensive simulation study in an attempt to compare their finite sample performance. It would seem from the results that variance estimators with adaptively chosen regularisation parameters perform admirably over a broad range of sparsity and signal strength settings. Finally, some intial theoretical analyses pertaining to these types of estimators are proposed and developed.
-
regression shrinkage and selection via the lasso
Journal of the royal statistical society series b-methodological, 1996Co-Authors: Robert TibshiraniAbstract:SumMARY We propose a new method for estimation in linear models. The 'lasso' minimizes the Residual Sum of squares subject to the Sum of the absolute value of the coefficients being less than a constant. Because of the nature of this constraint it tends to produce some coefficients that are exactly 0 and hence gives interpretable models. Our simulation studies suggest that the lasso enjoys some of the favourable properties of both subset selection and ridge regression. It produces interpretable models like subset selection and exhibits the stability of ridge regression. There is also an interesting relationship with recent work in adaptive function estimation by Donoho and Johnstone. The lasso idea is quite general and can be applied in a variety of statistical models: extensions to generalized regression models and tree-based models are briefly described.