The Experts below are selected from a list of 17919 Experts worldwide ranked by ideXlab platform

Sanghamitra Bandyopadhyay - One of the best experts on this subject based on the ideXlab platform.

  • Development of Some Line Symmetry Based Cluster Validity Indices
    2014 International Conference on Soft Computing and Machine Intelligence, 2014
    Co-Authors: Sudipta Acharya, Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Most of the existing literatures use Euclidean distance based Cluster Validity measures in order to identify correct number of Clusters for different datasets. It is a very important consideration for Clustering. Symmetry can be considered as an important attribute for data Clustering. It can be of two types, point symmetry and line symmetry. In this paper we have introduced a newly developed line symmetry based distance in the definitions of four well known Cluster Validity indices, namely Xie Beni(XB) index, PBM index, FCM index and PS index to identify proper partitioning and accurate number of Clusters from five artificially generated datasets. Initially in order to obtain different partitions an existing genetic Clustering technique which uses line symmetry property (GALS Clustering) is applied on datasets varying the number of Clusters. We have also provided a comparative study of our proposed line symmetry based Cluster Validity indices with their original versions which follow euclidean distance based computation. From the experimental results, it is revealed that most of the Validity indices which follow Line symmetry based distances, perform better than euclidean distance based original Validity indices.

  • Some connectivity based Cluster Validity indices
    Applied Soft Computing, 2012
    Co-Authors: Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Identification of the correct number of Clusters and the appropriate partitioning technique are some important considerations in Clustering where several Cluster Validity indices, primarily utilizing the Euclidean distance, have been used in the literature. In this paper a new measure of connectivity is incorporated in the definitions of seven Cluster Validity indices namely, DB-index, Dunn-index, Generalized Dunn-index, PS-index, I-index, XB-index and SV-index, thereby yielding seven new Cluster Validity indices which are able to automatically detect Clusters of any shape, size or convexity as long as they are well-separated. Here connectivity is measured using a novel approach following the concept of relative neighborhood graph. It is empirically established that incorporation of the property of connectivity significantly improves the capabilities of these indices in identifying the appropriate number of Clusters. The well-known Clustering techniques, single linkage Clustering technique and K-means Clustering technique are used as the underlying partitioning algorithms. Results on eight artificially generated and three real-life data sets show that connectivity based Dunn-index performs the best as compared to all the other six indices. Comparisons are made with the original versions of these seven Cluster Validity indices.

  • use of a fuzzy granulation degranulation criterion for assessing Cluster Validity
    Fuzzy Sets and Systems, 2011
    Co-Authors: Sanghamitra Bandyopadhyay, Sriparna Saha, Witold Pedrycz
    Abstract:

    The identification of a suitable Clustering algorithm to partition data and assessment of the Validity of the resultant partitioning are ongoing quests in unsupervised learning. In this study, a fuzzy granulation-degranulation criterion is proposed to evaluate the goodness of a fuzzy partitioning of the data. This, in turn, is used to determine the appropriate Clustering algorithm suitable for a particular data set. In general, the quality of a partitioning is measured by computing the variance within it, which is a measure of compactness of the obtained partitioning. Here a new error function, which reflects how well the computed Cluster centers represent the whole data set, is used as the goodness measure of the obtained partitioning. Thus a Clustering algorithm, providing a good set of Cluster centers which approximate well the whole data set, is considered to be the most suited. Thereafter this new fuzzy granulation-degranulation criterion is used to develop six new Cluster Validity indices. These indices mimic the definitions of the existing and well-known Cluster Validity indices, such as PBM-index, XB-index, PS-index, FS-index, K-index and SV-index, but use the new fuzzy granulation-degranulation based error function instead of Cluster compactness. In order to evaluate the effectiveness of the proposed error function in correctly identifying the appropriate Clustering algorithm for a particular data set, eight well-known Clustering algorithms, K-means, Fuzzy C-means, GAK-means (genetic algorithm based K-means algorithm), a newly developed genetic point symmetry based Clustering technique (GAPS-Clustering), Average Linkage Clustering algorithm, Expectation Maximization (EM) Clustering algorithm, Self-Organizing Map (SOM) and Spectral Clustering technique are evaluated on a set of six artificially generated and six real-life data sets. Results show that GAK-means is the most appropriate for most of the data sets used for the experiments. Thereafter the effectiveness of the proposed Cluster Validity indices in identifying the appropriate number of Clusters automatically from different data sets are shown for above mentioned 12 data sets. For the purpose of comparison, results obtained with the original versions of the proposed Cluster Validity indices and results obtained by a density based Clustering technique are also presented.

  • Performance Evaluation of Some Symmetry-Based Cluster Validity Indexes
    IEEE Transactions on Systems Man and Cybernetics Part C (Applications and Reviews), 2009
    Co-Authors: Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Identification of the correct number of Clusters is an important consideration in Clustering where several Cluster Validity indexes, primarily utilizing the Euclidean distance, have been used in the literature. The property of symmetry is observed in most Clustering solutions. In this paper, the symmetry versions of nine Cluster Validity indexes, namely, Davies-Bouldin index, Dunn index, generalized Dunn index, point symmetry (PS) index, I index, Xie-Beni index, FS index, K index, and SV index, are proposed. It is empirically established that incorporation of the property of symmetry significantly improves the capabilities of these indexes in identifying the appropriate number of Clusters. A recently developed PS-based genetic Clustering technique, GAPS Clustering, is used as the underlying partitioning algorithm. Results on six artificially generated and five real-life datasets show that symmetry-distance-based I index performs the best as compared to all the other eight indexes.

Yasunori Endo - One of the best experts on this subject based on the ideXlab platform.

  • Cluster Validity Measures Based Agglomerative Hierarchical Clustering for Network Data
    Journal of Advanced Computational Intelligence and Intelligent Informatics, 2019
    Co-Authors: Yukihiro Hamasuna, Ryo Ozaki, Shusuke Nakano, Yasunori Endo
    Abstract:

    The Louvain method is a method of agglomerative hierarchical Clustering (AHC) that uses modularity as the merging criterion. Modularity is an evaluation measure for network partitions. Cluster Validity measures are also used to evaluate Cluster partitions and to determine the optimal number of Clusters. Several Cluster Validity measures are constructed considering the geometric features of Clusters. These measures and modularity are considered to be the same concept in the viewpoint of evaluating Cluster partitions. In this paper, Cluster Validity measures based agglomerative hierarchical Clustering (CVAHC) is proposed as a novel Clustering method for network data. The Cluster Validity measures are used as a merging criterion and an evaluation measure for network data in the proposed method. Numerical experiments show that Dunn’s and Xie-Beni’s indices for network partitions are useful for network Clustering.

  • Cluster Validity Measures for Network Data
    Journal of Advanced Computational Intelligence and Intelligent Informatics, 2018
    Co-Authors: Yukihiro Hamasuna, Ryo Ozaki, Daiki Kobayashi, Yasunori Endo
    Abstract:

    Modularity is one of the evaluation measures for network partitions and is used as the merging criterion in the Louvain method. To construct useful Cluster Validity measures and Clustering methods for network data, network Cluster Validity measures are proposed based on the traditional indices. The effectiveness of the proposed measures are compared and applied to determine the optimal number of Clusters. The network Cluster partitions of various network data which are generated from the Polaris dataset are obtained byk-medoids with Dijkstra’s algorithm and evaluated by the proposed measures as well as the modularity. Our numerical experiments show that the Dunn’s index and the Xie-Beni’s index-based measures are effective for network partitions compared to other indices.

  • Two-Stage Clustering Based on Cluster Validity Measures
    Journal of Advanced Computational Intelligence and Intelligent Informatics, 2018
    Co-Authors: Yukihiro Hamasuna, Ryo Ozaki, Yasunori Endo
    Abstract:

    To handle a large-scale object, a two-stage Clustering method has been previously proposed. The method generates a large number of Clusters during the first stage and merges Clusters during the second stage. In this paper, a novel two-stage Clustering method is proposed by introducing Cluster Validity measures as the merging criterion during the second stage. The significant Cluster Validity measures used to evaluate Cluster partitions and determine the suitable number of Clusters act as the criteria for merging Clusters. The performance of the proposed method based on six typical indices is compared with eight artificial datasets. These experiments show that a trace of the fuzzy covariance matrixWtrand its kernelizationKWtrare quite effective when applying the proposed method, and obtain better results than the other indices.

  • SMC - Agglomerative hierarchical Clustering based on local optimization for Cluster Validity measures
    2017 IEEE International Conference on Systems Man and Cybernetics (SMC), 2017
    Co-Authors: Ryo Ozaki, Yukihiro Hamasuna, Yasunori Endo
    Abstract:

    Modularity is an evaluation measure for graph Clustering. Louvain method is constructed by local optimization for modularity and is bottom up method as well as agglomerative hierarchical Clustering. Cluster Validity measures are used to evaluate Cluster partitions as well as modularity. They are traditional evaluation measures in the field of Clustering. We propose a novel graph Clustering which is based on agglomerative hierarchical Clustering. The proposed method in this study is constructed by local optimization for Cluster Validity measures. The effectiveness of the proposed method is shown through numerical examples. Numerical examples show that the proposed method has different Clustering propety from Louvain method because of the feature of Cluster Validity measures.

  • a study on Cluster Validity measures for Clustering network data
    Soft Computing, 2017
    Co-Authors: Yukihiro Hamasuna, Ryo Ozaki, Yasunori Endo
    Abstract:

    Modularity is one of the evaluation measures for network data and used as the criterion of merging two Clusters in Louvain method. To construct useful Cluster Validity measures for network data, the effectiveness of eight conventional Cluster Validity measures are compared with Modularity. Cluster partitions of six artificial network datasets are obtained by k-medoids and evaluated by Cluster Validity measures including Modularity. Numerical experiments show that the Dunn's index is effective in conventional Cluster Validity measures than other indices.

Jen-chieh Chiang - One of the best experts on this subject based on the ideXlab platform.

  • A Cluster Validity Measure With Outlier Detection for Support Vector Clustering
    IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics), 2008
    Co-Authors: Jeen-shing Wang, Jen-chieh Chiang
    Abstract:

    This paper focuses on the development of an effective Cluster Validity measure with outlier detection and Cluster merging algorithms for support vector Clustering (SVC). Since SVC is a kernel-based Clustering approach, the parameter of kernel functions and the soft-margin constants in Lagrangian functions play a crucial role in the Clustering results. The major contribution of this paper is that our proposed Validity measure and algorithms are capable of identifying ideal parameters for SVC to reveal a suitable Cluster configuration for a given data set. A Validity measure, which is based on a ratio of Cluster compactness to separation with outlier detection and a Cluster-merging mechanism, has been developed to automatically determine ideal parameters for the kernel functions and soft-margin constants as well. With these parameters, the SVC algorithm is capable of identifying the optimal number of Clusters with compact and smooth arbitrary-shaped Cluster contours for the given data set and increasing robustness to outliers and noise. Several simulations, including artificial and benchmark data sets, have been conducted to demonstrate the effectiveness of the proposed Cluster Validity measure for the SVC algorithm.

  • SMC - Support Vector Clustering with a Novel Cluster Validity Method
    2006 IEEE International Conference on Systems Man and Cybernetics, 2006
    Co-Authors: Jen-chieh Chiang, Jeen-shing Wang
    Abstract:

    This paper presents a novel Cluster Validity method for the support vector Clustering (SVC) algorithm to identify an optimal Cluster configuration of a given data set. The SVC algorithm is a kernel-based Clustering approach that groups a data set into Clusters with irregular shapes. Without a priori knowledge of the data sets, a Validity measure based on a ratio of Cluster compactness to separation with outlier detection has been developed to automatically determine suitable parameters of the kernel functions and soft-margin constants as well. A novel Validity measure has been developed to find optimal Cluster configurations through an effective parameter searching algorithm. Computer simulations have been conducted on benchmark data sets to demonstrate the effectiveness of the proposed Cluster Validity method.

Sriparna Saha - One of the best experts on this subject based on the ideXlab platform.

  • Development of Some Line Symmetry Based Cluster Validity Indices
    2014 International Conference on Soft Computing and Machine Intelligence, 2014
    Co-Authors: Sudipta Acharya, Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Most of the existing literatures use Euclidean distance based Cluster Validity measures in order to identify correct number of Clusters for different datasets. It is a very important consideration for Clustering. Symmetry can be considered as an important attribute for data Clustering. It can be of two types, point symmetry and line symmetry. In this paper we have introduced a newly developed line symmetry based distance in the definitions of four well known Cluster Validity indices, namely Xie Beni(XB) index, PBM index, FCM index and PS index to identify proper partitioning and accurate number of Clusters from five artificially generated datasets. Initially in order to obtain different partitions an existing genetic Clustering technique which uses line symmetry property (GALS Clustering) is applied on datasets varying the number of Clusters. We have also provided a comparative study of our proposed line symmetry based Cluster Validity indices with their original versions which follow euclidean distance based computation. From the experimental results, it is revealed that most of the Validity indices which follow Line symmetry based distances, perform better than euclidean distance based original Validity indices.

  • Some connectivity based Cluster Validity indices
    Applied Soft Computing, 2012
    Co-Authors: Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Identification of the correct number of Clusters and the appropriate partitioning technique are some important considerations in Clustering where several Cluster Validity indices, primarily utilizing the Euclidean distance, have been used in the literature. In this paper a new measure of connectivity is incorporated in the definitions of seven Cluster Validity indices namely, DB-index, Dunn-index, Generalized Dunn-index, PS-index, I-index, XB-index and SV-index, thereby yielding seven new Cluster Validity indices which are able to automatically detect Clusters of any shape, size or convexity as long as they are well-separated. Here connectivity is measured using a novel approach following the concept of relative neighborhood graph. It is empirically established that incorporation of the property of connectivity significantly improves the capabilities of these indices in identifying the appropriate number of Clusters. The well-known Clustering techniques, single linkage Clustering technique and K-means Clustering technique are used as the underlying partitioning algorithms. Results on eight artificially generated and three real-life data sets show that connectivity based Dunn-index performs the best as compared to all the other six indices. Comparisons are made with the original versions of these seven Cluster Validity indices.

  • HIS - A min-max distance based external Cluster Validity index: MMI
    2012 12th International Conference on Hybrid Intelligent Systems (HIS), 2012
    Co-Authors: Abhay Kumar Alok, Sriparna Saha, Asif Ekbal
    Abstract:

    Evaluating a given Clustering result is a very difficult problem in real world. Cluster Validity indices are developed for this purpose. There are two different types of Cluster Validity indices available : External and Internal. External Cluster Validity indices utilize some supervised information and internal Cluster Validity indices utilize the intrinsic structure of the data. In this paper a new external Cluster Validity index, MMI has been implemented based on Max-Min distance among data points and prior information based on structure of the data. A new probabilistic approach has been implemented to find the correct correspondence between the true and obtained Clustering. Genetic K-means algorithm (GAK-means) and single linkage have been used as the underlying Clustering techniques. Results of the proposed index for identifying the appropriate number of Clusters is shown for five artificial and two real-life data sets. GAK-means and single linkage Clustering techniques are used as the underlying partitioning techniques with the number of Clusters varied over a range. The MMI index is then used to determine the appropriate number of Clusters. The performance of MMI is compared with existing external Cluster Validity indices, adjusted rand index (ARI) and rand index (RI). It works well for two class and multi class data sets.

  • use of a fuzzy granulation degranulation criterion for assessing Cluster Validity
    Fuzzy Sets and Systems, 2011
    Co-Authors: Sanghamitra Bandyopadhyay, Sriparna Saha, Witold Pedrycz
    Abstract:

    The identification of a suitable Clustering algorithm to partition data and assessment of the Validity of the resultant partitioning are ongoing quests in unsupervised learning. In this study, a fuzzy granulation-degranulation criterion is proposed to evaluate the goodness of a fuzzy partitioning of the data. This, in turn, is used to determine the appropriate Clustering algorithm suitable for a particular data set. In general, the quality of a partitioning is measured by computing the variance within it, which is a measure of compactness of the obtained partitioning. Here a new error function, which reflects how well the computed Cluster centers represent the whole data set, is used as the goodness measure of the obtained partitioning. Thus a Clustering algorithm, providing a good set of Cluster centers which approximate well the whole data set, is considered to be the most suited. Thereafter this new fuzzy granulation-degranulation criterion is used to develop six new Cluster Validity indices. These indices mimic the definitions of the existing and well-known Cluster Validity indices, such as PBM-index, XB-index, PS-index, FS-index, K-index and SV-index, but use the new fuzzy granulation-degranulation based error function instead of Cluster compactness. In order to evaluate the effectiveness of the proposed error function in correctly identifying the appropriate Clustering algorithm for a particular data set, eight well-known Clustering algorithms, K-means, Fuzzy C-means, GAK-means (genetic algorithm based K-means algorithm), a newly developed genetic point symmetry based Clustering technique (GAPS-Clustering), Average Linkage Clustering algorithm, Expectation Maximization (EM) Clustering algorithm, Self-Organizing Map (SOM) and Spectral Clustering technique are evaluated on a set of six artificially generated and six real-life data sets. Results show that GAK-means is the most appropriate for most of the data sets used for the experiments. Thereafter the effectiveness of the proposed Cluster Validity indices in identifying the appropriate number of Clusters automatically from different data sets are shown for above mentioned 12 data sets. For the purpose of comparison, results obtained with the original versions of the proposed Cluster Validity indices and results obtained by a density based Clustering technique are also presented.

  • Performance Evaluation of Some Symmetry-Based Cluster Validity Indexes
    IEEE Transactions on Systems Man and Cybernetics Part C (Applications and Reviews), 2009
    Co-Authors: Sriparna Saha, Sanghamitra Bandyopadhyay
    Abstract:

    Identification of the correct number of Clusters is an important consideration in Clustering where several Cluster Validity indexes, primarily utilizing the Euclidean distance, have been used in the literature. The property of symmetry is observed in most Clustering solutions. In this paper, the symmetry versions of nine Cluster Validity indexes, namely, Davies-Bouldin index, Dunn index, generalized Dunn index, point symmetry (PS) index, I index, Xie-Beni index, FS index, K index, and SV index, are proposed. It is empirically established that incorporation of the property of symmetry significantly improves the capabilities of these indexes in identifying the appropriate number of Clusters. A recently developed PS-based genetic Clustering technique, GAPS Clustering, is used as the underlying partitioning algorithm. Results on six artificially generated and five real-life datasets show that symmetry-distance-based I index performs the best as compared to all the other eight indexes.

Jeen-shing Wang - One of the best experts on this subject based on the ideXlab platform.

  • A Cluster Validity Measure With Outlier Detection for Support Vector Clustering
    IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics), 2008
    Co-Authors: Jeen-shing Wang, Jen-chieh Chiang
    Abstract:

    This paper focuses on the development of an effective Cluster Validity measure with outlier detection and Cluster merging algorithms for support vector Clustering (SVC). Since SVC is a kernel-based Clustering approach, the parameter of kernel functions and the soft-margin constants in Lagrangian functions play a crucial role in the Clustering results. The major contribution of this paper is that our proposed Validity measure and algorithms are capable of identifying ideal parameters for SVC to reveal a suitable Cluster configuration for a given data set. A Validity measure, which is based on a ratio of Cluster compactness to separation with outlier detection and a Cluster-merging mechanism, has been developed to automatically determine ideal parameters for the kernel functions and soft-margin constants as well. With these parameters, the SVC algorithm is capable of identifying the optimal number of Clusters with compact and smooth arbitrary-shaped Cluster contours for the given data set and increasing robustness to outliers and noise. Several simulations, including artificial and benchmark data sets, have been conducted to demonstrate the effectiveness of the proposed Cluster Validity measure for the SVC algorithm.

  • SMC - Support Vector Clustering with a Novel Cluster Validity Method
    2006 IEEE International Conference on Systems Man and Cybernetics, 2006
    Co-Authors: Jen-chieh Chiang, Jeen-shing Wang
    Abstract:

    This paper presents a novel Cluster Validity method for the support vector Clustering (SVC) algorithm to identify an optimal Cluster configuration of a given data set. The SVC algorithm is a kernel-based Clustering approach that groups a data set into Clusters with irregular shapes. Without a priori knowledge of the data sets, a Validity measure based on a ratio of Cluster compactness to separation with outlier detection has been developed to automatically determine suitable parameters of the kernel functions and soft-margin constants as well. A novel Validity measure has been developed to find optimal Cluster configurations through an effective parameter searching algorithm. Computer simulations have been conducted on benchmark data sets to demonstrate the effectiveness of the proposed Cluster Validity method.