The Experts below are selected from a list of 17262 Experts worldwide ranked by ideXlab platform
Michael I Jordan - One of the best experts on this subject based on the ideXlab platform.
-
matrix variate Dirichlet Process priors with applications
Bayesian Analysis, 2014Co-Authors: Zhihua Zhang, Guang Dai, Dakan Wang, Michael I JordanAbstract:In this paper we propose a matrix-variate Dirichlet Process (MATDP) for modeling the joint prior of a set of random matrices. Our approach is able to share statistical strength among regression coe cient matrices due to the clustering property of the Dirichlet Process. Moreover, since the base probability measure is de ned as a matrix-variate distribution, the dependence among the elements of each random matrix is described via the matrixvariate distribution. We apply MATDP to multivariate supervised learning problems. In particular, we devise a nonparametric discriminative model and a nonparametric latent factor model. The interest is in considering correlations both across response variables (or covariates) and across response vectors. We derive MCMC algorithms for posterior inference and prediction, and illustrate the application of the models to multivariate regression, multi-class classi cation and multi-label prediction problems.
-
small variance asymptotics for exponential family Dirichlet Process mixture models
Neural Information Processing Systems, 2012Co-Authors: Ke Jiang, Brian Kulis, Michael I JordanAbstract:Sampling and variational inference techniques are two standard methods for inference in probabilistic models, but for many problems, neither approach scales effectively to large-scale data. An alternative is to relax the probabilistic model into a non-probabilistic formulation which has a scalable associated algorithm. This can often be fulfilled by performing small-variance asymptotics, i.e., letting the variance of particular distributions in the model go to zero. For instance, in the context of clustering, such an approach yields connections between the k-means and EM algorithms. In this paper, we explore small-variance asymptotics for exponential family Dirichlet Process (DP) and hierarchical Dirichlet Process (HDP) mixture models. Utilizing connections between exponential family distributions and Bregman divergences, we derive novel clustering algorithms from the asymptotic limit of the DP and HDP mixtures that features the scalability of existing hard clustering methods as well as the flexibility of Bayesian nonparametric models. We focus on special cases of our analysis for discrete-data problems, including topic modeling, and we demonstrate the utility of our results by applying variants of our algorithms to problems arising in vision and document analysis.
-
matrix variate Dirichlet Process mixture models
International Conference on Artificial Intelligence and Statistics, 2010Co-Authors: Zhihua Zhang, Guang Dai, Michael I JordanAbstract:We are concerned with a multivariate response regression problem where the interest is in considering correlations both across response variates and across response samples. In this paper we develop a new Bayesian nonparametric model for such a setting based on Dirichlet Process priors. Building on an additive kernel model, we allow each sample to have its own regression matrix. Although this overcomplete representation could in principle sufier from severe overfltting problems, we are able to provide efiective control over the model via a matrix-variate Dirichlet Process prior on the regression matrices. Our model is able to share statistical strength among regression matrices due to the clustering property of the Dirichlet Process. We make use of a Markov chain Monte Carlo algorithm for inference and prediction. Compared with other Bayesian kernel models, our model has advantages in both computational and statistical e‐ciency.
-
AISTATS - Matrix-Variate Dirichlet Process Mixture Models
2010Co-Authors: Zhihua Zhang, Guang Dai, Michael I JordanAbstract:We are concerned with a multivariate response regression problem where the interest is in considering correlations both across response variates and across response samples. In this paper we develop a new Bayesian nonparametric model for such a setting based on Dirichlet Process priors. Building on an additive kernel model, we allow each sample to have its own regression matrix. Although this overcomplete representation could in principle sufier from severe overfltting problems, we are able to provide efiective control over the model via a matrix-variate Dirichlet Process prior on the regression matrices. Our model is able to share statistical strength among regression matrices due to the clustering property of the Dirichlet Process. We make use of a Markov chain Monte Carlo algorithm for inference and prediction. Compared with other Bayesian kernel models, our model has advantages in both computational and statistical e‐ciency.
-
Nonparametric empirical Bayes for the Dirichlet Process mixture model
Statistics and Computing, 2006Co-Authors: Jon D. Mcauliffe, David M. Blei, Michael I JordanAbstract:The Dirichlet Process prior allows flexible nonparametric mixture modeling. The number of mixture components is not specified in advance and can grow as new data arrive. However, analyses based on the Dirichlet Process prior are sensitive to the choice of the parameters, including an infinite-dimensional distributional parameter G _0. Most previous applications have either fixed G _0 as a member of a parametric family or treated G _0 in a Bayesian fashion, using parametric prior specifications. In contrast, we have developed an adaptive nonparametric method for constructing smooth estimates of G _0. We combine this method with a technique for estimating α, the other Dirichlet Process parameter, that is inspired by an existing characterization of its maximum-likelihood estimator. Together, these estimation procedures yield a flexible empirical Bayes treatment of Dirichlet Process mixtures. Such a treatment is useful in situations where smooth point estimates of G _0 are of intrinsic interest, or where the structure of G _0 cannot be conveniently modeled with the usual parametric prior families. Analysis of simulated and real-world datasets illustrates the robustness of this approach.
David M. Blei - One of the best experts on this subject based on the ideXlab platform.
-
a split merge mcmc algorithm for the hierarchical Dirichlet Process
arXiv: Machine Learning, 2012Co-Authors: Chong Wang, David M. BleiAbstract:The hierarchical Dirichlet Process (HDP) has become an important Bayesian nonparametric model for grouped data, such as document collections. The HDP is used to construct a flexible mixed-membership model where the number of components is determined by the data. As for most Bayesian nonparametric models, exact posterior inference is intractable---practitioners use Markov chain Monte Carlo (MCMC) or variational inference. Inspired by the split-merge MCMC algorithm for the Dirichlet Process (DP) mixture model, we describe a novel split-merge MCMC sampling algorithm for posterior inference in the HDP. We study its properties on both synthetic data and text corpora. We find that split-merge MCMC for the HDP can provide significant improvements over traditional Gibbs sampling, and we give some understanding of the data properties that give rise to larger improvements.
-
Dirichlet Process Mixtures of Generalized Linear Models
Journal of Machine Learning Research, 2011Co-Authors: Lauren A. Hannah, David M. Blei, Warren B. PowellAbstract:We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new class of methods for nonparametric regression. Given a data set of input-response pairs, the DP-GLM produces a global model of the joint distribution through a mixture of local generalized linear models. DP-GLMs allow both continuous and categorical inputs, and can model the same class of responses that can be modeled with a generalized linear model. We study the properties of the DP-GLM, and show why it provides better predictions and density estimates than existing Dirichlet Process mixture regression models. We give conditions for weak consistency of the joint distribution and pointwise consistency of the regression estimate.
-
Dirichlet Process mixtures of generalized linear models
arXiv: Machine Learning, 2009Co-Authors: Lauren A. Hannah, David M. Blei, Warren B. PowellAbstract:We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We prove conditions for the asymptotic unbiasedness of the DP-GLM regression mean function estimate. We also give examples for when those conditions hold, including models for compactly supported continuous distributions and a model with continuous covariates and categorical response. We empirically analyze the properties of the DP-GLM and why it provides better results than existing Dirichlet Process mixture regression models. We evaluate DP-GLM on several data sets, comparing it to modern methods of nonparametric regression like CART, Bayesian trees and Gaussian Processes. Compared to existing techniques, the DP-GLM provides a single model (and corresponding inference algorithms) that performs well in many regression settings.
-
Nonparametric empirical Bayes for the Dirichlet Process mixture model
Statistics and Computing, 2006Co-Authors: Jon D. Mcauliffe, David M. Blei, Michael I JordanAbstract:The Dirichlet Process prior allows flexible nonparametric mixture modeling. The number of mixture components is not specified in advance and can grow as new data arrive. However, analyses based on the Dirichlet Process prior are sensitive to the choice of the parameters, including an infinite-dimensional distributional parameter G _0. Most previous applications have either fixed G _0 as a member of a parametric family or treated G _0 in a Bayesian fashion, using parametric prior specifications. In contrast, we have developed an adaptive nonparametric method for constructing smooth estimates of G _0. We combine this method with a technique for estimating α, the other Dirichlet Process parameter, that is inspired by an existing characterization of its maximum-likelihood estimator. Together, these estimation procedures yield a flexible empirical Bayes treatment of Dirichlet Process mixtures. Such a treatment is useful in situations where smooth point estimates of G _0 are of intrinsic interest, or where the structure of G _0 cannot be conveniently modeled with the usual parametric prior families. Analysis of simulated and real-world datasets illustrates the robustness of this approach.
-
variational methods for the Dirichlet Process
International Conference on Machine Learning, 2004Co-Authors: David M. Blei, Michael I JordanAbstract:Variational inference methods, including mean field methods and loopy belief propagation, have been widely used for approximate probabilistic inference in graphical models. While often less accurate than MCMC, variational methods provide a fast deterministic approximation to marginal and conditional probabilities. Such approximations can be particularly useful in high dimensional problems where sampling methods are too slow to be effective. A limitation of current methods, however, is that they are restricted to parametric probabilistic models. MCMC does not have such a limitation; indeed, MCMC samplers have been developed for the Dirichlet Process (DP), a nonparametric distribution on distributions (Ferguson, 1973) that is the cornerstone of Bayesian nonparametric statistics (Escobar & West, 1995; Neal, 2000). In this paper, we develop a mean-field variational approach to approximate inference for the Dirichlet Process, where the approximate posterior is based on the truncated stick-breaking construction (Ishwaran & James, 2001). We compare our approach to DP samplers for Gaussian DP mixture models.
Luai Al Labadi - One of the best experts on this subject based on the ideXlab platform.
-
Goodness-of-fit tests based on the distance between the Dirichlet Process and its base measure
Journal of Nonparametric Statistics, 2014Co-Authors: Luai Al Labadi, Mahmoud ZarepourAbstract:The Dirichlet Process is a fundamental tool in studying Bayesian nonparametric inference. The Dirichlet Process has several sum representations, where each one of these representations highlights some aspects of this important Process. In this paper, we use the sum representations of the Dirichlet Process to derive explicit expressions that are used to calculate Kolmogorov, Levy, and Cramer–von Mises distances between the Dirichlet Process and its base measure. The derived expressions of the distance are used to select a proper value for the concentration parameter of the Dirichlet Process. These tools are also used in a goodness-of-fit test. Illustrative examples and simulation results are included.
-
On a rapid simulation of the Dirichlet Process
Statistics & Probability Letters, 2012Co-Authors: Mahmoud Zarepour, Luai Al LabadiAbstract:Abstract We describe a simple, yet efficient, procedure for approximating the Levy measure of a Gamma ( α , 1 ) random variable. We use this approximation to derive a finite sum-representation that converges almost surely to Ferguson’s representation of the Dirichlet Process. This approximation is written based on arrivals of a homogeneous Poisson Process. We compare the efficiency of our approximation to several other well-known approximations of the Dirichlet Process and demonstrate a significant improvement.
-
The Dirichlet Process with Large Concentration Parameter
arXiv: Statistics Theory, 2011Co-Authors: Luai Al Labadi, Mahmoud ZarepourAbstract:Ferguson's Dirichlet Process plays an important role in nonparametric Bayesian inference. Let $P_a$ be the Dirichlet Process in $\mathbb{R}$ with a base probability measure $H$ and a concentration parameter $a>0.$ In this paper, we show that $\sqrt {a} \big(P_a((-\infty,t]) -H((-\infty,t])\big)$ converges to a certain Brownian bridge as $a \to \infty.$ We also derive a certain Glivenko-Cantelli theorem for the Dirichlet Process. Using the functional delta method, the weak convergence of the quantile Process is also obtained. A large concentration parameter occurs when a statistician puts too much emphasize on his/her prior guess. This scenario also happens when the sample size is large and the posterior is used to make inference.
-
On a Rapid Simulation of the Dirichlet Process
arXiv: Machine Learning, 2011Co-Authors: Mahmoud Zarepour, Luai Al LabadiAbstract:We describe a simple and efficient procedure for approximating the L\'evy measure of a $\text{Gamma}(\alpha,1)$ random variable. We use this approximation to derive a finite sum-representation that converges almost surely to Ferguson's representation of the Dirichlet Process based on arrivals of a homogeneous Poisson Process. We compare the efficiency of our approximation to several other well known approximations of the Dirichlet Process and demonstrate a substantial improvement.
Alan E. Gelfand - One of the best experts on this subject based on the ideXlab platform.
-
A Partition Dirichlet Process Model for Functional Data Analysis
Sankhya B, 2020Co-Authors: Christoph Hellmayr, Alan E. GelfandAbstract:Recently, extensions of the Dirichlet Process to the functional domain have been presented in the literature. These Processes can be classified based on the type of labeling they induce across different arguments in the domain, i.e., marginal or joint labeling. We show that marginal labeling Processes have undesirable properties if functional observations are assumed to be almost surely continuous. Joint labeling Processes avoid this undesirable behavior. They are specified through finite dimensional distributions for locations across the entire domain. They range from a common label for all arguments (relatively easy to fit) to joint local label selection at every argument (computationally very demanding). Here, we offer a middle ground - a joint labeling Process which partitions the domain of the stochastic Process and assign labels to individual partition elements. We call the proposed model a partition functional Dirichlet Process and show that it can outperform both of the foregoing extremes of joint labeling. Given data that is a sample from each function in a collection of independent functions over the given domain, we employ this Process as a prior for the true functions. We show results from simulation studies as well as a dataset arising from reflectance curves to demonstrate the performance of this Process and make comparison with the two extreme labeling models.
-
the nested Dirichlet Process
Journal of the American Statistical Association, 2008Co-Authors: Abel Rodriguez, David B. Dunson, Alan E. GelfandAbstract:In multicenter studies, subjects in different centers may have different outcome distributions. This article is motivated by the problem of nonparametric modeling of these distributions, borrowing information across centers while also allowing centers to be clustered. Starting with a stick-breaking representation of the Dirichlet Process (DP), we replace the random atoms with random probability measures drawn from a DP. This results in a nested DP prior, which can be placed on the collection of distributions for the different centers, with centers drawn from the same DP component automatically clustered together. Theoretical properties are discussed, and an efficient Markov chain Monte Carlo algorithm is developed for computation. The methods are illustrated using a simulation study and an application to quality of care in U.S. hospitals.
-
Generalized spatial Dirichlet Process models
Biometrika, 2007Co-Authors: Jason A. Duan, Michele Guindani, Alan E. GelfandAbstract:SUMMARY Many models for the study of point-referenced data explicitly introduce spatial random effects to capture residual spatial association. These spatial effects are customarily modelled as a zeromean stationary Gaussian Process. The spatial Dirichlet Process introduced by Gelfand et al. (2005) produces a random spatial Process which is neither Gaussian nor stationary. Rather, it varies about a Process that is assumed to be stationary and Gaussian. The spatial Dirichlet Process arises as a probability-weighted collection of random surfaces. This can be limiting for modelling and inferential purposes since it insists that a Process realization must be one of these surfaces. We introduce a random distribution for the spatial effects that allows different surface selection at different sites. Moreover, we can specify the model so that the marginal distribution of the effect at each site still comes from a Dirichlet Process. The development is offered constructively, providing a multivariate extension of the stick-breaking representation of the weights. We then introduce mixing using this generalized spatial Dirichlet Process. We illustrate with a simulated dataset of independent replications and note that we can embed the generalized Process within a dynamic model specification to eliminate the independence assumption.
-
Bayesian Nonparametric Spatial Modeling With Dirichlet Process Mixing
Journal of the American Statistical Association, 2005Co-Authors: Alan E. Gelfand, Athanasios Kottas, Steven N MaceachernAbstract:Customary modeling for continuous point-referenced data assumes a Gaussian Process that is often taken to be stationary. When such models are fitted within a Bayesian framework, the unknown parameters of the Process are assumed to be random, so a random Gaussian Process results. Here we propose a novel spatial Dirichlet Process mixture model to produce a random spatial Process that is neither Gaussian nor stationary. We first develop a spatial Dirichlet Process model for spatial data and discuss its properties. Because of familiar limitations associated with direct use of Dirichlet Process models, we introduce mixing by convolving this Process with a pure error Process. We then examine properties of models created through such Dirichlet Process mixing. In the Bayesian framework, we implement posterior inference using Gibbs sampling. Spatial prediction raises interesting questions, but these can be handled. Finally, we illustrate the approach using simulated data, as well as a dataset involving precipitati...
Kenichi Kurihara - One of the best experts on this subject based on the ideXlab platform.
-
practical collapsed variational bayes inference for hierarchical Dirichlet Process
Knowledge Discovery and Data Mining, 2012Co-Authors: Issei Sato, Kenichi Kurihara, Hiroshi NakagawaAbstract:We propose a novel collapsed variational Bayes (CVB) inference for the hierarchical Dirichlet Process (HDP). While the existing CVB inference for the HDP variant of latent Dirichlet allocation (LDA) is more complicated and harder to implement than that for LDA, the proposed algorithm is simple to implement, does not require variance counts to be maintained, does not need to set hyper-parameters, and has good predictive performance.
-
collapsed variational Dirichlet Process mixture models
International Joint Conference on Artificial Intelligence, 2007Co-Authors: Kenichi Kurihara, Max Welling, Yee Whye TehAbstract:Nonparametric Bayesian mixture models, in particular Dirichlet Process (DP) mixture models, have shown great promise for density estimation and data clustering. Given the size of today's datasets, computational efficiency becomes an essential ingredient in the applicability of these techniques to real world data. We study and experimentally compare a number of variational Bayesian (VB) approximations to the DP mixture model. In particular we consider the standard VB approximation where parameters are assumed to be independent from cluster assignment variables, and a novel collapsed VB approximation where mixture weights are marginalized out. For both VB approximations we consider two different ways to approximate the DP, by truncating the stick-breaking construction, and by using a finite mixture model with a symmetric Dirichlet prior.
-
accelerated variational Dirichlet Process mixtures
Neural Information Processing Systems, 2006Co-Authors: Kenichi Kurihara, Max Welling, Nikos VlassisAbstract:Dirichlet Process (DP) mixture models are promising candidates for clustering applications where the number of clusters is unknown a priori. Due to computational considerations these models are unfortunately unsuitable for large scale data-mining applications. We propose a class of deterministic accelerated DP mixture models that can routinely handle millions of data-cases. The speedup is achieved by incorporating kd-trees into a variational Bayesian algorithm for DP mixtures in the stick-breaking representation, similar to that of Blei and Jordan (2005). Our algorithm differs in the use of kd-trees and in the way we handle truncation: we only assume that the variational distributions are fixed at their priors after a certain level. Experiments show that speedups relative to the standard variational algorithm can be significant.