The Experts below are selected from a list of 332805 Experts worldwide ranked by ideXlab platform

Peter C Austin - One of the best experts on this subject based on the ideXlab platform.

  • pisces did not have increased heart failure data driven comparisons of binary proportions between levels of a Categorical Variable can result in incorrect statistical significance levels
    Journal of Clinical Epidemiology, 2008
    Co-Authors: Peter C Austin, Meredith A Goldwasser
    Abstract:

    Abstract Objective We examined the impact on statistical inference when a χ2 test is used to compare the proportion of successes in the level of a Categorical Variable that has the highest observed proportion of successes with the proportion of successes in all other levels of the Categorical Variable combined. Study Design and Setting Monte Carlo simulations and a case study examining the association between astrological sign and hospitalization for heart failure. Results A standard χ2 test results in an inflation of the type I error rate, with the type I error rate increasing as the number of levels of the Categorical Variable increases. Using a standard χ2 test, the hospitalization rate for Pisces was statistically significantly different from that of the other 11 astrological signs combined (P = 0.026). After accounting for the fact that the selection of Pisces was based on it having the highest observed proportion of heart failure hospitalizations, subjects born under the sign of Pisces no longer had a significantly higher rate of heart failure hospitalization compared to the other residents of Ontario (P = 0.152). Conclusions Post hoc comparisons of the proportions of successes across different levels of a Categorical Variable can result in incorrect inferences.

  • inflation of the type i error rate when a continuous confounding Variable is categorized in logistic regression analyses
    Statistics in Medicine, 2004
    Co-Authors: Peter C Austin, Lawrence J Brunner
    Abstract:

    This paper demonstrates an inflation of the type I error rate that occurs when testing the statistical significance of a continuous risk factor after adjusting for a correlated continuous confounding Variable that has been divided into a Categorical Variable. We used Monte Carlo simulation methods to assess the inflation of the type I error rate when testing the statistical significance of a risk factor after adjusting for a continuous confounding Variable that has been divided into categories. We found that the inflation of the type I error rate increases with increasing sample size, as the correlation between the risk factor and the confounding Variable increases, and with a decrease in the number of categories into which the confounder is divided. Even when the confounder is divided in a five-level Categorical Variable, the inflation of the type I error rate remained high when both the sample size and the correlation between the risk factor and the confounder were high. Copyright 2004 John Wiley & Sons, Ltd.

  • Inflation of the type I error rate when a continuous confounding Variable is categorized in logistic regression analyses.
    Statistics in Medicine, 2004
    Co-Authors: Peter C Austin, Lawrence J Brunner
    Abstract:

    This paper demonstrates an inflation of the type I error rate that occurs when testing the statistical significance of a continuous risk factor after adjusting for a correlated continuous confounding Variable that has been divided into a Categorical Variable. We used Monte Carlo simulation methods to assess the inflation of the type I error rate when testing the statistical significance of a risk factor after adjusting for a continuous confounding Variable that has been divided into categories. We found that the inflation of the type I error rate increases with increasing sample size, as the correlation between the risk factor and the confounding Variable increases, and with a decrease in the number of categories into which the confounder is divided. Even when the confounder is divided in a five-level Categorical Variable, the inflation of the type I error rate remained high when both the sample size and the correlation between the risk factor and the confounder were high.

Lawrence J Brunner - One of the best experts on this subject based on the ideXlab platform.

  • inflation of the type i error rate when a continuous confounding Variable is categorized in logistic regression analyses
    Statistics in Medicine, 2004
    Co-Authors: Peter C Austin, Lawrence J Brunner
    Abstract:

    This paper demonstrates an inflation of the type I error rate that occurs when testing the statistical significance of a continuous risk factor after adjusting for a correlated continuous confounding Variable that has been divided into a Categorical Variable. We used Monte Carlo simulation methods to assess the inflation of the type I error rate when testing the statistical significance of a risk factor after adjusting for a continuous confounding Variable that has been divided into categories. We found that the inflation of the type I error rate increases with increasing sample size, as the correlation between the risk factor and the confounding Variable increases, and with a decrease in the number of categories into which the confounder is divided. Even when the confounder is divided in a five-level Categorical Variable, the inflation of the type I error rate remained high when both the sample size and the correlation between the risk factor and the confounder were high. Copyright 2004 John Wiley & Sons, Ltd.

  • Inflation of the type I error rate when a continuous confounding Variable is categorized in logistic regression analyses.
    Statistics in Medicine, 2004
    Co-Authors: Peter C Austin, Lawrence J Brunner
    Abstract:

    This paper demonstrates an inflation of the type I error rate that occurs when testing the statistical significance of a continuous risk factor after adjusting for a correlated continuous confounding Variable that has been divided into a Categorical Variable. We used Monte Carlo simulation methods to assess the inflation of the type I error rate when testing the statistical significance of a risk factor after adjusting for a continuous confounding Variable that has been divided into categories. We found that the inflation of the type I error rate increases with increasing sample size, as the correlation between the risk factor and the confounding Variable increases, and with a decrease in the number of categories into which the confounder is divided. Even when the confounder is divided in a five-level Categorical Variable, the inflation of the type I error rate remained high when both the sample size and the correlation between the risk factor and the confounder were high.

Clayton V Deutsch - One of the best experts on this subject based on the ideXlab platform.

  • Multiple Point Metrics to Assess Categorical Variable Models
    Natural Resources Research, 2010
    Co-Authors: Jeff B. Boisvert, Michael J. Pyrcz, Clayton V Deutsch
    Abstract:

    Geostatistical models should be checked to ensure consistency with conditioning data and statistical inputs. These are minimum acceptance criteria. Often the first and second-order statistics such as the histogram and variogram of simulated geological realizations are compared to the input parameters to check the reasonableness of the simulation implementation. Assessing the reproduction of statistics beyond second-order is often not considered because the “correct” higher order statistics are rarely known. With multiple point simulation (MPS) geostatistical methods, practitioners are now explicitly modeling higher-order statistics taken from a training image (TI). This article explores methods for extending minimum acceptance criteria to multiple point statistical comparisons between geostatistical realizations made with MPS algorithms and the associated TI. The intent is to assess how well the geostatistical models have reproduced the input statistics of the TI; akin to assessing the histogram and variogram reproduction in traditional semivariogram-based geostatistics. A number of metrics are presented to compare the input multiple point statistics of the TI with the statistics of the geostatistical realizations. These metrics are (1) first and second-order statistics, (2) trends, (3) the multiscale histogram, (4) the multiple point density function, and (5) the missing bins in the multiple point density function. A case study using MPS realizations is presented to demonstrate the proposed metrics; however, the metrics are not limited to specific MPS realizations. Comparisons could be made between any reference numerical analogue model and any simulated Categorical Variable model.

  • cleaning Categorical Variable lithofacies realizations with maximum a posteriori selection
    Computers & Geosciences, 1998
    Co-Authors: Clayton V Deutsch
    Abstract:

    Abstract Categorical Variable images frequently present unrealistic small scale variations (noise) that are an artefact of measurement error or the geostatistical simulation method. In addition to their unpleasing appearance, these variations may have an impact on subsequent petrophysical property modeling and flow simulation. Thus, there is a need to clean such realizations without altering their desirable features (reproduction of local data and large scale spatial structure). This paper presents a straightforward algorithm and program for such cleaning. Existing algorithms such as erosion/dilation, quantile-transformation, or iterative simulated annealing-based approaches for image cleaning have significant limitations when dealing with more than two categories. The key idea behind the proposed method is to retain at each location the most probable lithofacies type based on the surrounding lithofacies types, the proximity to conditioning data, and any mismatch from the global target proportion. A number of examples are presented. FORTRAN source code for the proposed algorithm, available at the IAMG web site, is documented.

Ivy Liu - One of the best experts on this subject based on the ideXlab platform.

  • strategies for modeling a Categorical Variable allowing multiple category choices
    Sociological Methods & Research, 2001
    Co-Authors: Alan Agresti, Ivy Liu
    Abstract:

    This article discusses strategies for modeling a Categorical Variable when subjects can select any subset of the categories. With c outcome categories, the models relate to a c-dimensional binary response, with each component indicating whether a particular category is chosen. The strategies are the following: (1) Using logit models directly for the marginal distribution of each component; this accounts for dependence among the component responses but does not treat the dependence as an integral part of the model. (2) Using logit models containing subject random effects to generate the dependence among the components; this approach is limited by implying nonnegative associations having a certain exchangeability. (3) Using loglinear modeling; quasi-symmetric ones are useful but are limited to estimation of within-subject effects. Marginal logit models less fully describe the dependence patterns for the data but require fewer assumptions and focus more directly on the effects of greatest substantive interest.

Zdenka Prokopova - One of the best experts on this subject based on the ideXlab platform.

  • Categorical Variable Segmentation Model for Software Development Effort Estimation
    IEEE Access, 2019
    Co-Authors: Petr Silhavy, Radek Silhavy, Zdenka Prokopova
    Abstract:

    This paper proposes a new software development effort estimation model. The new model's design is based on the function point analysis, Categorical Variable segmentation (CVS), and stepwise regression. The stepwise regression method is used for the creation of the unique estimation model of each segment. The estimation accuracy of the proposed model is compared to clustering-based models and the international function point user group model. It is shown that the proposed model increases estimation accuracy when compared to baseline methods: non-clustered functional point analysis and clustering-based models. The new CVS model achieves a significantly higher accuracy than the baseline methods.