The Experts below are selected from a list of 352500 Experts worldwide ranked by ideXlab platform

Masashi Sugiyama - One of the best experts on this subject based on the ideXlab platform.

  • online direct Density Ratio estimation applied to inlier based outlier detection
    Neural Computation, 2015
    Co-Authors: Marthinus Christoffel Du Plessis, Hiroaki Shiino, Masashi Sugiyama
    Abstract:

    Many machine learning problems, such as nonstationarity adaptation, outlier detection, dimensionality reduction, and conditional Density estimation, can be effectively solved by using the Ratio of probability densities. Since the naive two-step procedure of first estimating the probability densities and then taking their Ratio performs poorly, methods to directly estimate the Density Ratio from two sets of samples without Density estimation have been extensively studied recently. However, these methods are batch algorithms that use the whole data set to estimate the Density Ratio, and they are inefficient in the online setup, where training samples are provided sequentially and solutions are updated incrementally without storing previous samples. In this letter, we propose two online Density-Ratio estimators based on the adaptive regularization of weight vectors. Through experiments on inlier-based outlier detection, we demonstrate the usefulness of the proposed methods.

  • Density Ratio Hidden Markov Models
    arXiv: Machine Learning, 2013
    Co-Authors: John Quinn, Masashi Sugiyama
    Abstract:

    Hidden Markov models and their variants are the predominant sequential classification method in such domains as speech recognition, bioinformatics and natural language processing. Being generative rather than discriminative models, however, their classification performance is a drawback. In this paper we apply ideas from the field of Density Ratio estimation to bypass the difficult step of learning likelihood functions in HMMs. By reformulating inference and model fitting in terms of Density Ratios and applying a fast kernel-based estimation method, we show that it is possible to obtain a striking increase in discriminative performance while retaining the probabilistic qualities of the HMM. We demonstrate experimentally that this formulation makes more efficient use of training data than alternative approaches.

  • Relative Density-Ratio estimation for robust distribution comparison.
    Neural computation, 2013
    Co-Authors: Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
    Abstract:

    Divergence estimators based on direct approximation of Density Ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since Density-Ratio functions often possess high fluctuation, divergence estimation is a challenging task in practice. In this letter, we use relative divergences for distribution comparison, which involves approximation of relative Density Ratios. Since relative Density Ratios are always smoother than corresponding ordinary Density Ratios, our proposed method is favorable in terms of nonparametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.

  • statistical analysis of kernel based least squares Density Ratio estimation
    Machine Learning, 2012
    Co-Authors: Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
    Abstract:

    The Ratio of two probability densities can be used for solving various machine learning tasks such as covariate shift adaptation (importance sampling), outlier detection (likelihood-Ratio test), feature selection (mutual information), and conditional probability estimation. Several methods of directly estimating the Density Ratio have recently been developed, e.g., moment matching estimation, maximum-likelihood Density-Ratio estimation, and least-squares Density-Ratio fitting. In this paper, we propose a kernelized variant of the least-squares method for Density-Ratio estimation, which is called kernel unconstrained least-squares importance fitting (KuLSIF). We investigate its fundamental statistical properties including a non-parametric convergence rate, an analytic-form solution, and a leave-one-out cross-validation score. We further study its relation to other kernel-based Density-Ratio estimators. In experiments, we numerically compare various kernel-based Density-Ratio estimation methods, and show that KuLSIF compares favorably with other approaches.

  • Density Ratio estimation in machine learning
    2012
    Co-Authors: Masashi Sugiyama, Taiji Suzuki, Takafumi Kanamori
    Abstract:

    Machine learning is an interdisciplinary field of science and engineering that studies mathematical theories and practical applications of systems that learn. This book introduces theories, methods, and applications of Density Ratio estimation, which is a newly emerging paradigm in the machine learning community. Various machine learning problems such as non-stationarity adaptation, outlier detection, dimensionality reduction, independent component analysis, clustering, classification, and conditional Density estimation can be systematically solved via the estimation of probability Density Ratios. The authors offer a comprehensive introduction of various Density Ratio estimators including methods via Density estimation, moment matching, probabilistic classification, Density fitting, and Density Ratio fitting as well as describing how these can be applied to machine learning. The book also provides mathematical theories for Density Ratio estimation including parametric and non-parametric convergence analysis and numerical stability analysis to complete the first and definitive treatment of the entire framework of Density Ratio estimation in machine learning.

Taiji Suzuki - One of the best experts on this subject based on the ideXlab platform.

  • NIPS - Trimmed Density Ratio Estimation
    2017
    Co-Authors: Song Liu, Taiji Suzuki, Akiko Takeda, Kenji Fukumizu
    Abstract:

    Density Ratio estimation is a vital tool in both machine learning and statistical community. However, due to the unbounded nature of Density Ratio, the estimation proceudre can be vulnerable to corrupted data points, which often pushes the estimated Ratio toward infinity. In this paper, we present a robust estimator which automatically identifies and trims outliers. The proposed estimator has a convex formulation, and the global optimum can be obtained via subgradient descent. We analyze the parameter estimation error of this estimator under high-dimensional settings. Experiments are conducted to verify the effectiveness of the estimator.

  • Trimmed Density Ratio Estimation
    arXiv: Machine Learning, 2017
    Co-Authors: Song Liu, Taiji Suzuki, Akiko Takeda, Kenji Fukumizu
    Abstract:

    Density Ratio estimation is a vital tool in both machine learning and statistical community. However, due to the unbounded nature of Density Ratio, the estimation procedure can be vulnerable to corrupted data points, which often pushes the estimated Ratio toward infinity. In this paper, we present a robust estimator which automatically identifies and trims outliers. The proposed estimator has a convex formulation, and the global optimum can be obtained via subgradient descent. We analyze the parameter estimation error of this estimator under high-dimensional settings. Experiments are conducted to verify the effectiveness of the estimator.

  • Relative Density-Ratio estimation for robust distribution comparison.
    Neural computation, 2013
    Co-Authors: Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
    Abstract:

    Divergence estimators based on direct approximation of Density Ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since Density-Ratio functions often possess high fluctuation, divergence estimation is a challenging task in practice. In this letter, we use relative divergences for distribution comparison, which involves approximation of relative Density Ratios. Since relative Density Ratios are always smoother than corresponding ordinary Density Ratios, our proposed method is favorable in terms of nonparametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.

  • statistical analysis of kernel based least squares Density Ratio estimation
    Machine Learning, 2012
    Co-Authors: Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
    Abstract:

    The Ratio of two probability densities can be used for solving various machine learning tasks such as covariate shift adaptation (importance sampling), outlier detection (likelihood-Ratio test), feature selection (mutual information), and conditional probability estimation. Several methods of directly estimating the Density Ratio have recently been developed, e.g., moment matching estimation, maximum-likelihood Density-Ratio estimation, and least-squares Density-Ratio fitting. In this paper, we propose a kernelized variant of the least-squares method for Density-Ratio estimation, which is called kernel unconstrained least-squares importance fitting (KuLSIF). We investigate its fundamental statistical properties including a non-parametric convergence rate, an analytic-form solution, and a leave-one-out cross-validation score. We further study its relation to other kernel-based Density-Ratio estimators. In experiments, we numerically compare various kernel-based Density-Ratio estimation methods, and show that KuLSIF compares favorably with other approaches.

  • Density Ratio estimation in machine learning
    2012
    Co-Authors: Masashi Sugiyama, Taiji Suzuki, Takafumi Kanamori
    Abstract:

    Machine learning is an interdisciplinary field of science and engineering that studies mathematical theories and practical applications of systems that learn. This book introduces theories, methods, and applications of Density Ratio estimation, which is a newly emerging paradigm in the machine learning community. Various machine learning problems such as non-stationarity adaptation, outlier detection, dimensionality reduction, independent component analysis, clustering, classification, and conditional Density estimation can be systematically solved via the estimation of probability Density Ratios. The authors offer a comprehensive introduction of various Density Ratio estimators including methods via Density estimation, moment matching, probabilistic classification, Density fitting, and Density Ratio fitting as well as describing how these can be applied to machine learning. The book also provides mathematical theories for Density Ratio estimation including parametric and non-parametric convergence analysis and numerical stability analysis to complete the first and definitive treatment of the entire framework of Density Ratio estimation in machine learning.

Takafumi Kanamori - One of the best experts on this subject based on the ideXlab platform.

  • Relative Density-Ratio estimation for robust distribution comparison.
    Neural computation, 2013
    Co-Authors: Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
    Abstract:

    Divergence estimators based on direct approximation of Density Ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since Density-Ratio functions often possess high fluctuation, divergence estimation is a challenging task in practice. In this letter, we use relative divergences for distribution comparison, which involves approximation of relative Density Ratios. Since relative Density Ratios are always smoother than corresponding ordinary Density Ratios, our proposed method is favorable in terms of nonparametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.

  • statistical analysis of kernel based least squares Density Ratio estimation
    Machine Learning, 2012
    Co-Authors: Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
    Abstract:

    The Ratio of two probability densities can be used for solving various machine learning tasks such as covariate shift adaptation (importance sampling), outlier detection (likelihood-Ratio test), feature selection (mutual information), and conditional probability estimation. Several methods of directly estimating the Density Ratio have recently been developed, e.g., moment matching estimation, maximum-likelihood Density-Ratio estimation, and least-squares Density-Ratio fitting. In this paper, we propose a kernelized variant of the least-squares method for Density-Ratio estimation, which is called kernel unconstrained least-squares importance fitting (KuLSIF). We investigate its fundamental statistical properties including a non-parametric convergence rate, an analytic-form solution, and a leave-one-out cross-validation score. We further study its relation to other kernel-based Density-Ratio estimators. In experiments, we numerically compare various kernel-based Density-Ratio estimation methods, and show that KuLSIF compares favorably with other approaches.

  • Density Ratio estimation in machine learning
    2012
    Co-Authors: Masashi Sugiyama, Taiji Suzuki, Takafumi Kanamori
    Abstract:

    Machine learning is an interdisciplinary field of science and engineering that studies mathematical theories and practical applications of systems that learn. This book introduces theories, methods, and applications of Density Ratio estimation, which is a newly emerging paradigm in the machine learning community. Various machine learning problems such as non-stationarity adaptation, outlier detection, dimensionality reduction, independent component analysis, clustering, classification, and conditional Density estimation can be systematically solved via the estimation of probability Density Ratios. The authors offer a comprehensive introduction of various Density Ratio estimators including methods via Density estimation, moment matching, probabilistic classification, Density fitting, and Density Ratio fitting as well as describing how these can be applied to machine learning. The book also provides mathematical theories for Density Ratio estimation including parametric and non-parametric convergence analysis and numerical stability analysis to complete the first and definitive treatment of the entire framework of Density Ratio estimation in machine learning.

  • f divergence estimation and two sample homogeneity test under semiparametric Density Ratio models
    IEEE Transactions on Information Theory, 2012
    Co-Authors: Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
    Abstract:

    A Density Ratio is defined by the Ratio of two probability densities. We study the inference problem of Density Ratios and apply a semiparametric Density-Ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the f-divergence between two probability densities is estimated using a Density-Ratio estimator. The f -divergence estimator is then exploited for the two-sample homogeneity test. We derive an optimal estimator of f-divergence in the sense of the asymptotic variance in a semiparametric setting, and provide a statistic for two-sample homogeneity test based on the optimal estimator. We prove that the proposed test dominates the existing empirical likelihood score test. Through numerical studies, we illustrate the adequacy of the asymptotic theory for finite-sample inference.

  • Density-Ratio matching under the Bregman divergence: a unified framework of Density-Ratio estimation
    Annals of the Institute of Statistical Mathematics, 2011
    Co-Authors: Masashi Sugiyama, Taiji Suzuki, Takafumi Kanamori
    Abstract:

    Estimation of the Ratio of probability densities has attracted a great deal of attention since it can be used for addressing various statistical paradigms. A naive approach to Density-Ratio approximation is to first estimate numerator and denominator densities separately and then take their Ratio. However, this two-step approach does not perform well in practice, and methods for directly estimating Density Ratios without Density estimation have been explored. In this paper, we first give a comprehensive review of existing Density-Ratio estimation methods and discuss their pros and cons. Then we propose a new framework of Density-Ratio estimation in which a Density-Ratio model is fitted to the true Density-Ratio under the Bregman divergence. Our new framework includes existing approaches as special cases, and is substantially more general. Finally, we develop a robust Density-Ratio estimation method under the power divergence, which is a novel instance in our framework.

Konstantinos Fokianos - One of the best experts on this subject based on the ideXlab platform.

  • Safe Density Ratio modeling
    Statistics and Probability Letters, 2009
    Co-Authors: Kjell Konis, Konstantinos Fokianos
    Abstract:

    An important problem in logistic regression modeling is the existence of the maximum likelihood estimators. Especially when the sample size is small, the maximum likelihood estimator of the regression parameters does not exist if the data are completely, or quasi–completely separated. Recognizing that this phenomenon has a serious impact on the fitting of the Density Ratio model–which is a semiparametric model whose profile empirical log-likelihood has the logistic form because of the equivalence between prospective and retrospective sampling–we suggest a linear programming methodology for examining whether the maximum likelihood estimators of the finite dimensional parameter vector of the model exist. It is shown that the methodology can be effectively utilized in the analysis of case control gene expression data by identifying cases where the Density Ratio model cannot be applied. It is demonstrated that naive application of the Density Ratio model yields to erroneous conclusions.

  • Safe Density Ratio modeling
    Statistics & Probability Letters, 2009
    Co-Authors: Kjell Konis, Konstantinos Fokianos
    Abstract:

    An important problem in logistic regression modeling is the existence of the maximum likelihood estimators. In particular, when the sample size is small, the maximum likelihood estimator of the regression parameters does not exist if the data are completely, or quasicompletely separated. Recognizing that this phenomenon has a serious impact on the fitting of the Density Ratio model–which is a semiparametric model whose profile empirical log-likelihood has the logistic form because of the equivalence between prospective and retrospective sampling–we suggest a linear programming methodology for examining whether the maximum likelihood estimators of the finite dimensional parameter vector of the model exist. It is shown that the methodology can be effectively utilized in the analysis of case–control gene expression data by identifying cases where the Density Ratio model cannot be applied. It is demonstrated that naive application of the Density Ratio model yields erroneous conclusions.

  • Density Ratio Model Selection
    Journal of Statistical Computation and Simulation, 2007
    Co-Authors: Konstantinos Fokianos
    Abstract:

    The Density Ratio model presumes that the log-likelihood Ratio of two unknown densities is of some known parametric linear form. However, the choice of the functional form has an impact on both estimation and testing. The problem of over/underfitting in the context of the Density Ratio model is examined and the theory shows that bias and loss of efficiency are introduced when the model is misspecified. The problem of identifying the appropriate functional form for an application of the Density Ratio model is addressed by means of model selection criteria, which perform reasonably well. Several simulations integrate the presentation.

  • Inference for the Relative Treatment Effect with the Density Ratio Model
    Statistical Modelling, 2007
    Co-Authors: Konstantinos Fokianos, James F. Troendle
    Abstract:

    Consider the problem of estimating and testing the relative treatment effect between two populations based on a random sample from each distribution. Under the well established normal theory, inference is based on analysis of variance methods. However, there are many examples of skewed data which show that normal theory is not applicable. Then the problem of inference regarding the treatment effect can be attacked by standard nonparametric methods. In this note, we propose a semiparametric model, the so‐called Density Ratio model which specifies that the log‐likelihood Ratio of two densities is linear in some parameters. For testing hypotheses regarding the relative treatment effect, a robust test is obtained by employing the Density Ratio model for a suitable Box‐Cox transformation of the data. The transformation along with the Density Ratio model are estimated by maximum empirical likelihood. The new test procedure is studied theoretically and it is applied to real and simulated data. It is further compared with some nonparametric competitors, and it is found to have relatively high power across a wide variety of distributions including those outside the Density Ratio family.

Makoto Yamada - One of the best experts on this subject based on the ideXlab platform.

  • Relative Density-Ratio estimation for robust distribution comparison.
    Neural computation, 2013
    Co-Authors: Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
    Abstract:

    Divergence estimators based on direct approximation of Density Ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since Density-Ratio functions often possess high fluctuation, divergence estimation is a challenging task in practice. In this letter, we use relative divergences for distribution comparison, which involves approximation of relative Density Ratios. Since relative Density Ratios are always smoother than corresponding ordinary Density Ratios, our proposed method is favorable in terms of nonparametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.

  • relative Density Ratio estimation for robust distribution comparison
    arXiv: Machine Learning, 2011
    Co-Authors: Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
    Abstract:

    Divergence estimators based on direct approximation of Density-Ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since Density-Ratio functions often possess high fluctuation, divergence estimation is still a challenging task in practice. In this paper, we propose to use relative divergences for distribution comparison, which involves approximation of relative Density-Ratios. Since relative Density-Ratios are always smoother than corresponding ordinary Density-Ratios, our proposed method is favorable in terms of the non-parametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.