The Experts below are selected from a list of 46245 Experts worldwide ranked by ideXlab platform

Bernhard Scholkopf - One of the best experts on this subject based on the ideXlab platform.

  • on estimation of functional causal models post nonlinear causal model as an example
    International Conference on Data Mining, 2013
    Co-Authors: Kun Zhang, Zhikun Wang, Bernhard Scholkopf
    Abstract:

    Compared to constraint-based causal discovery, causal discovery based on functional causal models is able to identify the whole causal model under appropriate assumptions. Functional causal models represent the effect as a function of the direct causes together with an independent noise term. Examples include the linear non-Gaussian a cyclic model (LiNGAM), nonlinear additive noise model, and post-nonlinear (PNL) model. Currently there are two ways to estimate the parameters in the models, one is by dependence minimization, and the other is maximum likelihood. In this paper, we show that for any a cyclic functional causal model, minimizing the mutual information between the hypothetical cause and the noise term is equivalent to maximizing the data likelihood with a flexible model for the distribution of the noise term. We then focus on estimation of the PNL causal model, and propose to estimate it with the warped Gaussian process with the noise modeled by the mixture of Gaussians. As a Bayesian nonparametric approach, it outperforms the previous one based on mutual information minimization with nonlinear functions represented by multilayer perceptrons, we also show that unlike the Ordinary Regression, estimation results of the PNL causal model are sensitive to the assumption on the noise distribution. Experimental results on both synthetic and real data support our theoretical claims.

  • on causal discovery with cyclic additive noise models
    Neural Information Processing Systems, 2011
    Co-Authors: Joris M Mooij, Dominik Janzing, Tom Heskes, Bernhard Scholkopf
    Abstract:

    We study a particular class of cyclic causal models, where each variable is a (possibly nonlinear) function of its parents and additive noise. We prove that the causal graph of such models is generically identifiable in the bivariate, Gaussian-noise case. We also propose a method to learn such models from observational data. In the acyclic case, the method reduces to Ordinary Regression, but in the more challenging cyclic case, an additional term arises in the loss function, which makes it a special case of nonlinear independent component analysis. We illustrate the proposed method on synthetic data.

Norihiro Kato - One of the best experts on this subject based on the ideXlab platform.

  • nonlinear ridge Regression improves cell type specific differential expression analysis
    BMC Bioinformatics, 2021
    Co-Authors: Fumihiko Takeuchi, Norihiro Kato
    Abstract:

    Background Epigenome-wide association studies (EWAS) and differential gene expression analyses are generally performed on tissue samples, which consist of multiple cell types. Cell-type-specific effects of a trait, such as disease, on the omics expression are of interest but difficult or costly to measure experimentally. By measuring omics data for the bulk tissue, cell type composition of a sample can be inferred statistically. Subsequently, cell-type-specific effects are estimated by linear Regression that includes terms representing the interaction between the cell type proportions and the trait. This approach involves two issues, scaling and multicollinearity. Results First, although cell composition is analyzed in linear scale, differential methylation/expression is analyzed suitably in the logit/log scale. To simultaneously analyze two scales, we applied nonlinear Regression. Second, we show that the interaction terms are highly collinear, which is obstructive to Ordinary Regression. To cope with the multicollinearity, we applied ridge regularization. In simulated data, nonlinear ridge Regression attained well-balanced sensitivity, specificity and precision. Marginal model attained the lowest precision and highest sensitivity and was the only algorithm to detect weak signal in real data. Conclusion Nonlinear ridge Regression performed cell-type-specific association test on bulk omics data with well-balanced performance. The omicwas package for R implements nonlinear ridge Regression for cell-type-specific EWAS, differential gene expression and QTL analyses. The software is freely available from https://github.com/fumi-github/omicwas.

  • nonlinear ridge Regression improves robustness of cell type specific differential expression analysis
    bioRxiv, 2020
    Co-Authors: Fumihiko Takeuchi, Norihiro Kato
    Abstract:

    Background: Epigenome-wide association studies (EWAS) and differential gene expression analyses are generally performed on tissue samples, which consist of multiple cell types. Cell-type-specific effects of a trait, such as disease, on the omics expression are of interest but difficult or costly to measure experimentally. By measuring omics data for the bulk tissue, cell type composition of a sample can be inferred statistically. Subsequently, cell-type-specific effects are estimated by linear Regression that includes terms representing the interaction between the cell type proportions and the trait. This approach involves two issues, scaling and multicollinearity. Results: First, although cell composition is analyzed in linear scale, differential methylation/expression is analyzed suitably in the logit/log scale. To simultaneously analyze two scales, we applied nonlinear Regression. Second, we show that the interaction terms are highly collinear, which is obstructive to Ordinary Regression. To cope with the multicollinearity, we applied ridge regularization. In simulated data, nonlinear ridge Regression attained well-balanced sensitivity, specificity and precision. In real data, nonlinear ridge Regression detected signals consistently over the examined cases. Conclusion: Nonlinear ridge Regression performed cell-type-specific association test on bulk omics data more robustly than previous methods. The omicwas package for R implements nonlinear ridge Regression for cell-type-specific EWAS, differential gene expression and QTL analyses. The software is freely available from https://github.com/fumi-github/omicwas

Ross L Prentice - One of the best experts on this subject based on the ideXlab platform.

  • a risk set calibration method for failure time Regression by using a covariate reliability sample
    Journal of The Royal Statistical Society Series B-statistical Methodology, 2001
    Co-Authors: Sharon X Xie, C Y Wang, Ross L Prentice
    Abstract:

    Regression parameter estimation in the Cox failure time model is considered when Regression variables are subject to measurement error. Assuming that repeat Regression vector measurements adhere to a classical measurement model, we can consider an Ordinary Regression calibration approach in which the unobserved covariates are replaced by an estimate of their conditional expectation given available covariate measurements. However, since the rate of withdrawal from the risk set across the time axis, due to failure or censoring, will typically depend on covariates, we may improve the Regression parameter estimator by recalibrating within each risk set. The asymptotic and small sample properties of such a risk set Regression calibration estimator are studied. A simple estimator based on a least squares calibration in each risk set appears able to eliminate much of the bias that attends the Ordinary Regression calibration estimator under extreme measurement error circumstances. Corresponding asymptotic distribution theory is developed, small sample properties are studied using computer simulations and an illustration is provided.

Angela M Wood - One of the best experts on this subject based on the ideXlab platform.

  • the use of repeated blood pressure measures for cardiovascular risk prediction a comparison of statistical models in the aric study
    Statistics in Medicine, 2017
    Co-Authors: Michael J Sweeting, Jessica K Barrett, Simon G Thompson, Angela M Wood
    Abstract:

    : Many prediction models have been developed for the risk assessment and the prevention of cardiovascular disease in primary care. Recent efforts have focused on improving the accuracy of these prediction models by adding novel biomarkers to a common set of baseline risk predictors. Few have considered incorporating repeated measures of the common risk predictors. Through application to the Atherosclerosis Risk in Communities study and simulations, we compare models that use simple summary measures of the repeat information on systolic blood pressure, such as (i) baseline only; (ii) last observation carried forward; and (iii) cumulative mean, against more complex methods that model the repeat information using (iv) Ordinary Regression calibration; (v) risk-set Regression calibration; and (vi) joint longitudinal and survival models. In comparison with the baseline-only model, we observed modest improvements in discrimination and calibration using the cumulative mean of systolic blood pressure, but little further improvement from any of the complex methods. © 2016 The Authors. Statistics in Medicine Published by John Wiley & Sons Ltd.

Nicholas T Longford - One of the best experts on this subject based on the ideXlab platform.

  • an alternative to model selection in Ordinary Regression
    Statistics and Computing, 2003
    Co-Authors: Nicholas T Longford
    Abstract:

    The weaknesses of established model selection procedures based on hypothesis testing and similar criteria are discussed and an alternative based on synthetic (composite) estimation is proposed. It is developed for the problem of prediction in Ordinary Regression and its properties are explored by simulations for the simple Regression. Extensions to a general setting are described and an example with multiple Regression is analysed. Arguments are presented against using a selected model for any inferences.

  • random coefficient models
    1994
    Co-Authors: Nicholas T Longford
    Abstract:

    Quantitative social research relies heavily on data that originate either from surveys with sampling designs that depart from simple random sampling, or from observational studies with no formal sampling design. Simple random sampling is often not feasible, or its use would yield data with less information about certain features of interest, and it is often economically prohibitive. For example, in studies of school effectiveness it may be difficult to secure the cooperation of a school or a classroom. Therefore it would be rather wasteful to collect data from a small number of students in such a classroom. Data from a larger proportion, or from all the students, could be collected at a small additional expense, thus reducing the number of classrooms required for a sample to contain sufficient information for the intended purposes. Similarly, in household surveys, having contacted a selected individual, it would make sense to collect data from the rest of the members of the household at the same time. When this is done, we usually end up with data for which the standard assumptions of independence (such as in Ordinary Regression) are inappropriate.