The Experts below are selected from a list of 116331 Experts worldwide ranked by ideXlab platform
Ujjwal Maulik - One of the best experts on this subject based on the ideXlab platform.
-
Ensemble based rough fuzzy clustering for Categorical Data
Knowledge-Based Systems, 2015Co-Authors: Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal MaulikAbstract:Categorical Data is different from continuous Data, where the values of attribute do not follow any natural ordering. Moreover, inherent complexities like uncertainty, vagueness and overlapping among clusters make the analysis of real life Categorical Data set more difficult. Recent literature review shows that the well-known Categorical Data clustering techniques are using different similarity/dissimilarity measures to tackle the inherent complexities of the Categorical attribute values. Generally, it is hard to find single method and cluster validity measure that can be used as perfect or standard for all kinds of Categorical Data sets. Hence, in this paper first, a clustering method for Categorical Data is proposed by fusing rough set and fuzzy set theories. Subsequently, an ensemble based framework is designed with the recently proposed similarity/dissimilarity measures in order to have better clustering results for different types of Categorical Data sets. For this purpose, the proposed rough fuzzy clustering method is used sequentially with the integration of different measures to evolve the clustering solutions. Using consensus of these solutions, pure classified, semi rough and pure rough points are identified. Thereafter, machine learning method, called Random Forest, is used in incremental way to classify the semi and pure rough points using pure classified points to yield better clustering results. The performance of the proposed method has been demonstrated in comparison with several other recently developed clustering methods. Additionally, the selection of Random Forest in the proposed framework is justified by comparing its performance with other well-known machine learning methods like K-Nearest Neighbor and Support Vector Machine. Ten Categorical Data sets are used for the experimental purpose. Finally, statistical significance test has been conducted to judge the superiority of the results.
-
Differential fuzzy clustering for Categorical Data
2009 Proceeding of International Conference on Methods and Models in Computer Science (ICM2CS), 2009Co-Authors: Indrajit Saha, Ujjwal Maulik, NilanjanAbstract:Differential evolution has emerged as one of the fast, robust, and efficient global search heuristics of current interest. Besides its good convergence properties and suitability for parallelization, Differential evolution's main assets are its conceptual simplicity and ease of use. This paper describes an application of differential evolution to the fuzzy clustering for Categorical Data sets. The performance of the proposed method has been compared with the simulated annealing based fuzzy c-medoids clustering algorithm, fuzzy c-medoids, fuzzy c-modes and average linkage hierarchical clustering algorithm for two artificial and two real life Categorical Data sets. Statistical significance test has been carried out to establish the statistical significance of the proposed method.
-
IEEE Congress on Evolutionary Computation - Multiobjective approach to Categorical Data clustering
2007 IEEE Congress on Evolutionary Computation, 2007Co-Authors: Anirban Mukhopadhyay, Ujjwal MaulikAbstract:Categorical Data clustering has been gaining significant attention from researchers since the last few years, because most of the real life Data sets are Categorical in nature. In contrast to numerical domain, no natural ordering can be found among the elements of a Categorical domain. Hence no inherent distance measure, like the Euclidean distance, would work to compute the distance between two Categorical objects. Most of the clustering algorithms designed for Categorical Data are based on optimizing a single objective function. However, a single objective function is often not applicable for different kinds of Categorical Data sets. Motivated by this fact, in this article, the Categorical Data clustering problem has been modeled as a multiobjective optimization problem. A popular multiobjective genetic algorithm has been used in this regard to optimize two objectives simultaneously, thus generating a set of non-dominated solutions. The performance of the proposed algorithm has been compared with that of different well known Categorical Data clustering algorithms and demonstrated for a variety of synthetic and real life Categorical Data sets. Also a statistical significance test has been performed to establish the superiority of the proposed algorithm.
Chuangyin Dang - One of the best experts on this subject based on the ideXlab platform.
-
Space Structure and Clustering of Categorical Data
IEEE transactions on neural networks and learning systems, 2015Co-Authors: Yuhua Qian, Jiye Liang, Bing Liu, Chuangyin DangAbstract:Learning from Categorical Data plays a fundamental role in such areas as pattern recognition, machine learning, Data mining, and knowledge discovery. To effectively discover the group structure inherent in a set of Categorical objects, many Categorical clustering algorithms have been developed in the literature, among which $k$ -modes-type algorithms are very representative because of their good performance. Nevertheless, there is still much room for improving their clustering performance in comparison with the clustering algorithms for the numeric Data. This may arise from the fact that the Categorical Data lack a clear space structure as that of the numeric Data. To address this issue, we propose, in this paper, a novel Data-representation scheme for the Categorical Data, which maps a set of Categorical objects into a Euclidean space. Based on the Data-representation scheme, a general framework for space structure based Categorical clustering algorithms (SBC) is designed. This framework together with the applications of two kinds of dissimilarities leads two versions of the SBC-type algorithms. To verify the performance of the SBC-type algorithms, we employ as references four representative algorithms of the $k$ -modes-type algorithms. Experiments show that the proposed SBC-type algorithms significantly outperform the $k$ -modes-type algorithms.
-
a cluster centers initialization method for clustering Categorical Data
Expert Systems With Applications, 2012Co-Authors: Jiye Liang, Chuangyin DangAbstract:The leading partitional clustering technique, k-modes, is one of the most computationally efficient clustering methods for Categorical Data. However, the performance of the k-modes clustering algorithm which converges to numerous local minima strongly depends on initial cluster centers. Currently, most methods of initialization cluster centers are mainly for numerical Data. Due to lack of geometry for the Categorical Data, these methods used in cluster centers initialization for numerical Data are not applicable to Categorical Data. This paper proposes a novel initialization method for Categorical Data which is implemented to the k-modes algorithm. The method integrates the distance and the density together to select initial cluster centers and overcomes shortcomings of the existing initialization methods for Categorical Data. Experimental results illustrate the proposed initialization method is effective and can be applied to large Data sets for its linear time complexity with respect to the number of Data objects.
Bruno Scarpa - One of the best experts on this subject based on the ideXlab platform.
-
Bayesian inference on group differences in multivariate Categorical Data
Computational Statistics & Data Analysis, 2018Co-Authors: Massimiliano Russo, Daniele Durante, Bruno ScarpaAbstract:Abstract Multivariate Categorical Data are common in many fields. An illustrative example is provided by election polls studies assessing evidence of changes in voters’ opinions with their candidates preferences in the 2016 United States Presidential primaries or caucuses. Similar goals arise in routine applications, but current literature lacks a general methodology which combines flexibility, efficiency, and tractability in testing for group differences in multivariate Categorical Data at different – potentially complex – scales. This contribution addresses such goal by leveraging a Bayesian representation, which factorizes the joint probability mass function for the group variable and the multivariate Categorical Data as the product of the marginal probabilities for the groups and the conditional probability mass function of the multivariate Categorical Data, given the group membership. To enhance flexibility, the conditional probability mass function of the multivariate Categorical Data is defined via a group-dependent mixture of tensor factorizations which facilitates dimensionality reduction and borrowing of information, while providing tractable procedures for computation, and accurate tests assessing global and local group differences. The proposed methods are compared with popular competitors, and the improved performance is outlined in simulations and in American election polls studies.
-
Bayesian inference on group differences in multivariate Categorical Data
arXiv: Methodology, 2016Co-Authors: Massimiliano Russo, Daniele Durante, Bruno ScarpaAbstract:Multivariate Categorical Data are common in many fields. We are motivated by election polls studies assessing evidence of changes in voters opinions with their candidates preferences in the 2016 United States Presidential primaries or caucuses. Similar goals arise routinely in several applications, but current literature lacks a general methodology which combines flexibility, efficiency, and tractability in testing for group differences in multivariate Categorical Data at different---potentially complex---scales. We address this goal by leveraging a Bayesian representation which factorizes the joint probability mass function for the group variable and the multivariate Categorical Data as the product of the marginal probabilities for the groups, and the conditional probability mass function of the multivariate Categorical Data, given the group membership. To enhance flexibility, we define the conditional probability mass function of the multivariate Categorical Data via a group-dependent mixture of tensor factorizations, thus facilitating dimensionality reduction and borrowing of information, while providing tractable procedures for computation, and accurate tests assessing global and local group differences. We compare our methods with popular competitors, and discuss improved performance in simulations and in American election polls studies.
Indrajit Saha - One of the best experts on this subject based on the ideXlab platform.
-
Ensemble based rough fuzzy clustering for Categorical Data
Knowledge-Based Systems, 2015Co-Authors: Indrajit Saha, Jnanendra Prasad Sarkar, Ujjwal MaulikAbstract:Categorical Data is different from continuous Data, where the values of attribute do not follow any natural ordering. Moreover, inherent complexities like uncertainty, vagueness and overlapping among clusters make the analysis of real life Categorical Data set more difficult. Recent literature review shows that the well-known Categorical Data clustering techniques are using different similarity/dissimilarity measures to tackle the inherent complexities of the Categorical attribute values. Generally, it is hard to find single method and cluster validity measure that can be used as perfect or standard for all kinds of Categorical Data sets. Hence, in this paper first, a clustering method for Categorical Data is proposed by fusing rough set and fuzzy set theories. Subsequently, an ensemble based framework is designed with the recently proposed similarity/dissimilarity measures in order to have better clustering results for different types of Categorical Data sets. For this purpose, the proposed rough fuzzy clustering method is used sequentially with the integration of different measures to evolve the clustering solutions. Using consensus of these solutions, pure classified, semi rough and pure rough points are identified. Thereafter, machine learning method, called Random Forest, is used in incremental way to classify the semi and pure rough points using pure classified points to yield better clustering results. The performance of the proposed method has been demonstrated in comparison with several other recently developed clustering methods. Additionally, the selection of Random Forest in the proposed framework is justified by comparing its performance with other well-known machine learning methods like K-Nearest Neighbor and Support Vector Machine. Ten Categorical Data sets are used for the experimental purpose. Finally, statistical significance test has been conducted to judge the superiority of the results.
-
Differential fuzzy clustering for Categorical Data
2009 Proceeding of International Conference on Methods and Models in Computer Science (ICM2CS), 2009Co-Authors: Indrajit Saha, Ujjwal Maulik, NilanjanAbstract:Differential evolution has emerged as one of the fast, robust, and efficient global search heuristics of current interest. Besides its good convergence properties and suitability for parallelization, Differential evolution's main assets are its conceptual simplicity and ease of use. This paper describes an application of differential evolution to the fuzzy clustering for Categorical Data sets. The performance of the proposed method has been compared with the simulated annealing based fuzzy c-medoids clustering algorithm, fuzzy c-medoids, fuzzy c-modes and average linkage hierarchical clustering algorithm for two artificial and two real life Categorical Data sets. Statistical significance test has been carried out to establish the statistical significance of the proposed method.
Zhexue Huang - One of the best experts on this subject based on the ideXlab platform.
-
A fuzzy k-modes algorithm for clustering Categorical Data
IEEE Transactions on Fuzzy Systems, 1999Co-Authors: Zhexue HuangAbstract:This correspondence describes extensions to the fuzzy k-means algorithm for clustering Categorical Data. By using a simple matching dissimilarity measure for Categorical objects and modes instead of means for clusters, a new approach is developed, which allows the use of the k-means paradigm to efficiently cluster large Categorical Data sets. A fuzzy k-modes algorithm is presented and the effectiveness of the algorithm is demonstrated with experimental results.