The Experts below are selected from a list of 13869 Experts worldwide ranked by ideXlab platform

Shui Feng - One of the best experts on this subject based on the ideXlab platform.

Nizar Bouguila - One of the best experts on this subject based on the ideXlab platform.

  • predicting defect prone software modules using shifted scaled Dirichlet Distribution
    International Conference on Artificial Intelligence, 2018
    Co-Authors: Rua Alsuroji, Nizar Bouguila, Nuha Zamzami
    Abstract:

    Effective prediction of defect-prone software modules enables software developers to avoid the expensive costs in resources and efforts they might expense, and focus efficiently on quality assurance activities. Different classification methods have been applied previously to categorize a module in a system into two classes; defective or non-defective. Among the successful approaches, finite mixture modeling has been efficiently applied for solving this problem. This paper proposes the shifted-scaled Dirichlet model (SSDM) and evaluates its capability in predicting defect-prone software modules in the context of four NASA datasets. The results indicate that the prediction performance of SSDM is competitive to some previously used generative models.

  • unsupervised learning of finite mixtures using scaled Dirichlet Distribution and its application to software modules categorization
    International Conference on Industrial Technology, 2017
    Co-Authors: Bromensele Samuel Oboh, Nizar Bouguila
    Abstract:

    We have designed and implemented an unsupervised learning algorithm for finite mixture model using the scaled Dirichlet Distribution for multivariate proportional data. In this paper, the task of learning finite mixture model involves estimation of model parameters as well as inferring the hidden class information of our observed data. We made use of the expectation maximization algorithm to find the maximum likelihood estimate of our model parameters. This work, aims to address the flexibility challenge of the Dirichlet Distribution by introducing a Distribution that adds to it a scale parameter. This is important, because there is growing need for models that can fully describe the intrinsic nature of datasets. In addition, we applied our learning algorithm to synthetic datasets as well as to address the challenge of detecting fault prone software modules. Our proposed algorithm, makes it possible to discover these fault prone modules by harnessing their complexity-based attribute information. Finally, we compare our proposed model classification results with those from the Gaussian and Dirichlet mixture models.

  • dimensionality reduction of proportional data through data separation using Dirichlet Distribution
    International Conference on Image Analysis and Recognition, 2015
    Co-Authors: Walid Masoudimansour, Nizar Bouguila
    Abstract:

    In this paper, a novel method is proposed for dimensionality reduction of proportional data. Non-negative, unit-sum data, namely, proportional data emerges in many applications such as document classification, image classification using visual bag of words, etc. The introduced method is supervised and can be used for classification of data into binary classes. In the proposed method, the intra-class correlation is maximized while minimizing the interclass correlation, using a linear transform. Design of this transform is formulated as an optimization problem with proper cost function. The projected data is matched to two Dirichlet Distributions with careful parameter selection which allows to separate the classes in the Dirichlet parameter space. Finally, simulations are performed to demonstrate the effectiveness of the algorithm.

  • A Dirichlet Process Mixture of Generalized Dirichlet Distributions for Proportional Data Modeling
    Neural Networks, IEEE Transactions on, 2010
    Co-Authors: Nizar Bouguila, Djemel Ziou
    Abstract:

    In this paper, we propose a clustering algorithm based on both Dirichlet processes and generalized Dirichlet Distribution which has been shown to be very flexible for proportional data modeling. Our approach can be viewed as an extension of the finite generalized Dirichlet mixture model to the infinite case. The extension is based on nonparametric Bayesian analysis. This clustering algorithm does not require the specification of the number of mixture components to be given in advance and estimates it in a principled manner. Our approach is Bayesian and relies on the estimation of the posterior Distribution of clusterings using Gibbs sampler. Through some applications involving real-data classification and image databases categorization using visual words, we show that clustering via infinite mixture models offers a more powerful and robust performance than classic finite mixtures.

  • high dimensional unsupervised selection and estimation of a finite generalized Dirichlet mixture model based on minimum message length
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2007
    Co-Authors: Nizar Bouguila, Djemel Ziou
    Abstract:

    We consider the problem of determining the structure of high-dimensional data without prior knowledge of the number of clusters. Data are represented by a finite mixture model based on the generalized Dirichlet Distribution. The generalized Dirichlet Distribution has a more general covariance structure than the Dirichlet Distribution and offers high flexibility and ease of use for the approximation of both symmetric and asymmetric Distributions. This makes the generalized Dirichlet Distribution more practical and useful. An important problem in mixture modeling is the determination of the number of clusters. Indeed, a mixture with too many or too few components may not be appropriate to approximate the true model. Here, we consider the application of the minimum message length (MML) principle to determine the number of clusters. The MML is derived so as to choose the number of clusters in the mixture model that best describes the data. A comparison with other selection criteria is performed. The validation involves synthetic data, real data clustering, and two interesting real applications: classification of Web pages, and texture database summarization for efficient retrieval.

Djemel Ziou - One of the best experts on this subject based on the ideXlab platform.

Yuki Kawakubo - One of the best experts on this subject based on the ideXlab platform.

  • latent mixture modeling for clustered data
    Statistics and Computing, 2019
    Co-Authors: Shonosuke Sugasawa, Genya Kobayashi, Yuki Kawakubo
    Abstract:

    This article proposes a mixture modeling approach to estimating cluster-wise conditional Distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent Distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet Distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation–Maximization algorithm. We also extend the model such that the Distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as mixed effects models through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.

  • latent mixture modeling for clustered data
    arXiv: Methodology, 2017
    Co-Authors: Shonosuke Sugasawa, Genya Kobayashi, Yuki Kawakubo
    Abstract:

    This article proposes a mixture modeling approach to estimating cluster-wise conditional Distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent Distributions, and propose a model in which each cluster-wise density is represented as a mixture of latent experts with cluster-wise mixing proportions distributed as Dirichlet Distribution. The model parameters are estimated by maximizing the marginal likelihood function using a newly developed Monte Carlo Expectation-Maximization algorithm. We also extend the model such that the Distribution of cluster-wise mixing proportions depends on some cluster-level covariates. The finite sample performance of the proposed model is compared with some existing mixture modeling approaches as well as linear mixed model through the simulation studies. The proposed model is also illustrated with the posted land price data in Japan.

Nathan Ross - One of the best experts on this subject based on the ideXlab platform.

  • stein s method for the poisson Dirichlet Distribution and the ewens sampling formula with applications to wright fisher models
    Annals of Applied Probability, 2021
    Co-Authors: Han L Gan, Nathan Ross
    Abstract:

    We provide a general theorem bounding the error in the approximation of a random measure of interest—for example, the empirical population measure of types in a Wright–Fisher model—and a Dirichlet process, which is a measure having Poisson–Dirichlet distributed atoms with i.i.d. labels from a diffuse Distribution. The implicit metric of the approximation theorem captures the sizes and locations of the masses, and so also yields bounds on the approximation between the masses of the measure of interest and the Poisson–Dirichlet Distribution. We apply the result to bound the error in the approximation of the stationary Distribution of types in the finite Wright–Fisher model with infinite-alleles mutation structure (not necessarily parent independent) by the Poisson–Dirichlet Distribution. An important consequence of our result is an explicit upper bound on the total variation distance between the random partition generated by sampling from a finite Wright–Fisher stationary Distribution, and the Ewens sampling formula. The bound is small if the sample size n is much smaller than N1/6log(N)−1/2, where N is the total population size. Our analysis requires a result of separate interest, giving an explicit bound on the second moment of the number of types of a finite Wright–Fisher stationary Distribution. The general approximation result follows from a new development of Stein’s method for the Dirichlet process, which follows by viewing the Dirichlet process as the stationary Distribution of a Fleming–Viot process, and then applying Barbour’s generator approach.