The Experts below are selected from a list of 13869 Experts worldwide ranked by ideXlab platform
Shui Feng - One of the best experts on this subject based on the ideXlab platform.
-
asymptotic behaviour of poisson Dirichlet Distribution and random energy model
2015Co-Authors: Shui Feng, Youzhou ZhouAbstract:The family of Poisson-Dirichlet Distributions is a collection of two-parameter probability Distributions \(\{PD(\alpha,\theta ): 0 \leq \alpha 0\}\) defined on the infinite-dimensional simplex. The parameters α and θ correspond to the stable and gamma component respectively. The Distribution PD(α, 0) arises in the thermodynamic limit of the Gibbs measure of Derrida’s Random Energy Model (REM) in the low temperature regime. In this setting α can be written as the ratio between the temperature T and a critical temperature T c . In this paper, we study the asymptotic behaviour of PD(α, θ) as α converges to one or equivalently when the temperature approaches the critical value T c .
-
asymptotic results for the two parameter poisson Dirichlet Distribution
Stochastic Processes and their Applications, 2010Co-Authors: Shui Feng, Fuqing GaoAbstract:Abstract The two-parameter Poisson–Dirichlet Distribution is the law of a sequence of decreasing nonnegative random variables with total sum one. It can be constructed from stable and gamma subordinators with the two parameters, α and θ , corresponding to the stable component and the gamma component respectively. The moderate deviation principle is established for the Distribution when θ approaches infinity, and the large deviation principle is established when both α and θ approach zero.
-
the poisson Dirichlet Distribution and related topics models and asymptotic behaviors
2010Co-Authors: Shui FengAbstract:Models.- The Poisson-Dirichlet Distribution.- The Two-Parameter Poisson-Dirichlet Distribution.- The Coalescent.- Stochastic Dynamics.- Particle Representation.- Asymptotic Behaviors.- Fluctuation Theorems.- Large Deviations for the Poisson-Dirichlet Distribution.- Large Deviations for the Dirichlet Processes.
-
Some Diffusion Processes Associated With Two Parameter Poisson-Dirichlet Distribution and Dirichlet Process
Probability Theory and Related Fields, 2009Co-Authors: Shui Feng, Wei SunAbstract:The two parameter Poisson–Dirichlet Distribution PD(α, θ) is the Distribution of an infinite dimensional random discrete probability. It is a generalization of Kingman’s Poisson–Dirichlet Distribution. The two parameter Dirichlet process \({\Pi_{\alpha,\theta,\nu_0}}\) is the law of a pure atomic random measure with masses following the two parameter Poisson–Dirichlet Distribution. In this article we focus on the construction and the properties of the infinite dimensional symmetric diffusion processes with respective symmetric measures PD(α, θ) and \({\Pi_{\alpha,\theta,\nu_0}}\). The methods used come from the theory of Dirichlet forms.
-
asymptotic results for the two parameter poisson Dirichlet Distribution
arXiv: Probability, 2009Co-Authors: Shui Feng, Fuqing GaoAbstract:The two-parameter Poisson-Dirichlet Distribution is the law of a sequence of decreasing nonnegative random variables with total sum one. It can be constructed from stable and Gamma subordinators with the two-parameters, $\alpha$ and $\theta$, corresponding to the stable component and Gamma component respectively. The moderate deviation principles are established for the two-parameter Poisson-Dirichlet Distribution and the corresponding homozygosity when $\theta$ approaches infinity, and the large deviation principle is established for the two-parameter Poisson-Dirichlet Distribution when both $\alpha$ and $\theta$ approach zero.
Nizar Bouguila - One of the best experts on this subject based on the ideXlab platform.
-
predicting defect prone software modules using shifted scaled Dirichlet Distribution
International Conference on Artificial Intelligence, 2018Co-Authors: Rua Alsuroji, Nizar Bouguila, Nuha ZamzamiAbstract:Effective prediction of defect-prone software modules enables software developers to avoid the expensive costs in resources and efforts they might expense, and focus efficiently on quality assurance activities. Different classification methods have been applied previously to categorize a module in a system into two classes; defective or non-defective. Among the successful approaches, finite mixture modeling has been efficiently applied for solving this problem. This paper proposes the shifted-scaled Dirichlet model (SSDM) and evaluates its capability in predicting defect-prone software modules in the context of four NASA datasets. The results indicate that the prediction performance of SSDM is competitive to some previously used generative models.
-
unsupervised learning of finite mixtures using scaled Dirichlet Distribution and its application to software modules categorization
International Conference on Industrial Technology, 2017Co-Authors: Bromensele Samuel Oboh, Nizar BouguilaAbstract:We have designed and implemented an unsupervised learning algorithm for finite mixture model using the scaled Dirichlet Distribution for multivariate proportional data. In this paper, the task of learning finite mixture model involves estimation of model parameters as well as inferring the hidden class information of our observed data. We made use of the expectation maximization algorithm to find the maximum likelihood estimate of our model parameters. This work, aims to address the flexibility challenge of the Dirichlet Distribution by introducing a Distribution that adds to it a scale parameter. This is important, because there is growing need for models that can fully describe the intrinsic nature of datasets. In addition, we applied our learning algorithm to synthetic datasets as well as to address the challenge of detecting fault prone software modules. Our proposed algorithm, makes it possible to discover these fault prone modules by harnessing their complexity-based attribute information. Finally, we compare our proposed model classification results with those from the Gaussian and Dirichlet mixture models.
-
dimensionality reduction of proportional data through data separation using Dirichlet Distribution
International Conference on Image Analysis and Recognition, 2015Co-Authors: Walid Masoudimansour, Nizar BouguilaAbstract:In this paper, a novel method is proposed for dimensionality reduction of proportional data. Non-negative, unit-sum data, namely, proportional data emerges in many applications such as document classification, image classification using visual bag of words, etc. The introduced method is supervised and can be used for classification of data into binary classes. In the proposed method, the intra-class correlation is maximized while minimizing the interclass correlation, using a linear transform. Design of this transform is formulated as an optimization problem with proper cost function. The projected data is matched to two Dirichlet Distributions with careful parameter selection which allows to separate the classes in the Dirichlet parameter space. Finally, simulations are performed to demonstrate the effectiveness of the algorithm.
-
A Dirichlet Process Mixture of Generalized Dirichlet Distributions for Proportional Data Modeling
Neural Networks, IEEE Transactions on, 2010Co-Authors: Nizar Bouguila, Djemel ZiouAbstract:In this paper, we propose a clustering algorithm based on both Dirichlet processes and generalized Dirichlet Distribution which has been shown to be very flexible for proportional data modeling. Our approach can be viewed as an extension of the finite generalized Dirichlet mixture model to the infinite case. The extension is based on nonparametric Bayesian analysis. This clustering algorithm does not require the specification of the number of mixture components to be given in advance and estimates it in a principled manner. Our approach is Bayesian and relies on the estimation of the posterior Distribution of clusterings using Gibbs sampler. Through some applications involving real-data classification and image databases categorization using visual words, we show that clustering via infinite mixture models offers a more powerful and robust performance than classic finite mixtures.
-
high dimensional unsupervised selection and estimation of a finite generalized Dirichlet mixture model based on minimum message length
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2007Co-Authors: Nizar Bouguila, Djemel ZiouAbstract:We consider the problem of determining the structure of high-dimensional data without prior knowledge of the number of clusters. Data are represented by a finite mixture model based on the generalized Dirichlet Distribution. The generalized Dirichlet Distribution has a more general covariance structure than the Dirichlet Distribution and offers high flexibility and ease of use for the approximation of both symmetric and asymmetric Distributions. This makes the generalized Dirichlet Distribution more practical and useful. An important problem in mixture modeling is the determination of the number of clusters. Indeed, a mixture with too many or too few components may not be appropriate to approximate the true model. Here, we consider the application of the minimum message length (MML) principle to determine the number of clusters. The MML is derived so as to choose the number of clusters in the mixture model that best describes the data. A comparison with other selection criteria is performed. The validation involves synthetic data, real data clustering, and two interesting real applications: classification of Web pages, and texture database summarization for efficient retrieval.
Djemel Ziou - One of the best experts on this subject based on the ideXlab platform.
-
A Dirichlet Process Mixture of Generalized Dirichlet Distributions for Proportional Data Modeling
Neural Networks, IEEE Transactions on, 2010Co-Authors: Nizar Bouguila, Djemel ZiouAbstract:In this paper, we propose a clustering algorithm based on both Dirichlet processes and generalized Dirichlet Distribution which has been shown to be very flexible for proportional data modeling. Our approach can be viewed as an extension of the finite generalized Dirichlet mixture model to the infinite case. The extension is based on nonparametric Bayesian analysis. This clustering algorithm does not require the specification of the number of mixture components to be given in advance and estimates it in a principled manner. Our approach is Bayesian and relies on the estimation of the posterior Distribution of clusterings using Gibbs sampler. Through some applications involving real-data classification and image databases categorization using visual words, we show that clustering via infinite mixture models offers a more powerful and robust performance than classic finite mixtures.
-
high dimensional unsupervised selection and estimation of a finite generalized Dirichlet mixture model based on minimum message length
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2007Co-Authors: Nizar Bouguila, Djemel ZiouAbstract:We consider the problem of determining the structure of high-dimensional data without prior knowledge of the number of clusters. Data are represented by a finite mixture model based on the generalized Dirichlet Distribution. The generalized Dirichlet Distribution has a more general covariance structure than the Dirichlet Distribution and offers high flexibility and ease of use for the approximation of both symmetric and asymmetric Distributions. This makes the generalized Dirichlet Distribution more practical and useful. An important problem in mixture modeling is the determination of the number of clusters. Indeed, a mixture with too many or too few components may not be appropriate to approximate the true model. Here, we consider the application of the minimum message length (MML) principle to determine the number of clusters. The MML is derived so as to choose the number of clusters in the mixture model that best describes the data. A comparison with other selection criteria is performed. The validation involves synthetic data, real data clustering, and two interesting real applications: classification of Web pages, and texture database summarization for efficient retrieval.
-
a powerful finite mixture model based on the generalized Dirichlet Distribution unsupervised learning and applications
International Conference on Pattern Recognition, 2004Co-Authors: Nizar Bouguila, Djemel ZiouAbstract:This paper presents a new finite mixture model based on a generalization of the Dirichlet Distribution. For the estimation of the parameters of this mixture we use a GEM (generalized expectation maximization) algorithm based on a Newton-Raphson step. The experimental results involve the comparison of the performance of Gaussian and generalized Dirichlet mixtures in the classification of several pattern-recognition data sets.
-
Unsupervised learning of a finite mixture model based on the Dirichlet Distribution and its application
Image Processing, IEEE Transactions on, 2004Co-Authors: Nizar Bouguila, Djemel Ziou, Jean VaillancourtAbstract:This paper presents an unsupervised algorithm for learning a finite mixture model from multivariate data. This mixture model is based on the Dirichlet Distribution, which offers high flexibility for modeling data. The proposed approach for estimating the parameters of a Dirichlet mixture is based on the maximum likelihood (ML) and Fisher scoring methods. Experimental results are presented for the following applications: estimation of artificial histograms, summarization of image databases for efficient retrieval, and human skin color modeling and its application to skin detection in multimedia databases.
-
novel mixtures based on the Dirichlet Distribution application to data and image classification
Machine Learning and Data Mining in Pattern Recognition, 2003Co-Authors: Nizar Bouguila, Djemel Ziou, Jean VaillancourtAbstract:The Dirichlet Distribution offers high flexibility for modeling data. This paper describes two new mixtures based on this density: the GDD (Generalized Dirichlet Distribution) and the MDD (Multinomial Dirichlet Distribution) mixtures. These mixtures will be used to model continuous and discrete data, respectively. We propose a method for estimating the parameters of these mixtures. The performance of our method is tested by contextual evaluations. In these evaluations we compare the performance of Gaussian and GDD mixtures in the classification of several pattern-recognition data sets and we apply the MDD mixture to the problem of summarizing image databases.
Yuki Kawakubo - One of the best experts on this subject based on the ideXlab platform.
-
latent mixture modeling for clustered data
Statistics and Computing, 2019Co-Authors: Shonosuke Sugasawa, Genya Kobayashi, Yuki KawakuboAbstract:This article proposes a mixture modeling approach to estimating cluster-wise conditional Distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent Distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet Distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation–Maximization algorithm. We also extend the model such that the Distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as mixed effects models through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.
-
latent mixture modeling for clustered data
arXiv: Methodology, 2017Co-Authors: Shonosuke Sugasawa, Genya Kobayashi, Yuki KawakuboAbstract:This article proposes a mixture modeling approach to estimating cluster-wise conditional Distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent Distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet Distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation-Maximization algorithm. We also extend the model such that the Distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as linear mixed model through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.
Nathan Ross - One of the best experts on this subject based on the ideXlab platform.
-
stein s method for the poisson Dirichlet Distribution and the ewens sampling formula with applications to wright fisher models
Annals of Applied Probability, 2021Co-Authors: Han L Gan, Nathan RossAbstract:We provide a general theorem bounding the error in the approximation of a random measure of interest—for example, the empirical population measure of types in a Wright–Fisher model—and a Dirichlet process, which is a measure having Poisson–Dirichlet distributed atoms with i.i.d. labels from a diffuse Distribution. The implicit metric of the approximation theorem captures the sizes and locations of the masses, and so also yields bounds on the approximation between the masses of the measure of interest and the Poisson–Dirichlet Distribution. We apply the result to bound the error in the approximation of the stationary Distribution of types in the finite Wright–Fisher model with infinite-alleles mutation structure (not necessarily parent independent) by the Poisson–Dirichlet Distribution. An important consequence of our result is an explicit upper bound on the total variation distance between the random partition generated by sampling from a finite Wright–Fisher stationary Distribution, and the Ewens sampling formula. The bound is small if the sample size n is much smaller than N1/6log(N)−1/2, where N is the total population size. Our analysis requires a result of separate interest, giving an explicit bound on the second moment of the number of types of a finite Wright–Fisher stationary Distribution. The general approximation result follows from a new development of Stein’s method for the Dirichlet process, which follows by viewing the Dirichlet process as the stationary Distribution of a Fleming–Viot process, and then applying Barbour’s generator approach.