The Experts below are selected from a list of 426 Experts worldwide ranked by ideXlab platform
Asoke K Nandi - One of the best experts on this subject based on the ideXlab platform.
-
Effect of Explicit Evaluation on Neural Connectivity Related to Listening to Unfamiliar Music
Frontiers in Human Neuroscience, 2017Co-Authors: Elvira Brattico, Basel Abu-jamous, Carlos S. Pereira, Thomas Jacobsen, Asoke K NandiAbstract:People can experience different emotions when listening to music. A growing number of studies have investigated the brain structures and neural connectivities associated with perceived emotions. However, very little is known about the effect of an explicit act of judgement on the neural processing of emotionally-valenced music. In this study, we adopted the novel Consensus clustering paradigm, called binarisation of Consensus Partition matrices (Bi-CoPaM), to study whether and how the conscious aesthetic evaluation of the music would modulate brain connectivity networks related to emotion and reward processing. Participants listened to music under three conditions – one involving a non-evaluative judgment, one involving an explicit evaluative aesthetic judgment, and one involving no judgement at all (passive listening only). During non-evaluative attentive listening we obtained auditory-limbic connectivity whereas when participants were asked to decide explicitly whether they liked or disliked the music excerpt, only two clusters of intercommunicating brain regions were found: one including areas related to auditory processing and action observation, and the other comprising higher-order structures involved with visual processing. Results indicate that explicit evaluative judgment has an impact on the neural auditory-limbic connectivity during affective processing of music.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets. Results Here, we propose an unsupervised method for the unification of clustering results from multiple datasets using external specifications (UNCLES) . This method has the ability to identify the subsets of genes consistently co-expressed in a subset of datasets while being poorly co-expressed in another subset of datasets, and to identify the subsets of genes consistently co-expressed in all given datasets. We also propose the M-N scatter plots validation technique and adopt it to set the parameters of UNCLES, such as the number of clusters, automatically. Additionally, we propose an approach for the synthesis of gene expression datasets using real data profiles in a way which combines the ground-truth-knowledge of synthetic data and the realistic expression values of real data, and therefore overcomes the problem of faithfulness of synthetic expression data modelling. By application to those datasets, we validate UNCLES while comparing it with other conventional clustering methods, and of particular relevance, biclustering methods. We further validate UNCLES by application to a set of 14 real genome-wide yeast datasets as it produces focused clusters that conform well to known biological facts. Furthermore, in-silico -based hypotheses regarding the function of a few previously unknown genes in those focused clusters are drawn. Conclusions The UNCLES method, the M-N scatter plots technique, and the expression data synthesis approach will have wide application for the comprehensive analysis of genomic and other sources of multiple complex biological datasets. Moreover, the derived in-silico -based biological hypotheses represent subjects for future functional studies.
-
Application of the Bi-CoPaM Method to Five Escherichia Coli Datasets Generated under Various Biological Conditions
Journal of Signal Processing Systems, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The increasing amounts of high-throughput biological datasets stimulate the information engineering and machine learning research community to direct more studies towards designing and applying novel methods which are sophisticated and specialised to tackle the problems that are specific in such datasets. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method tackles the problem of scrutinising multiple gene expression microarray datasets to identify the subsets of genes which are consistently co-expressed across them. It allows for clustering results which better reflect the biological fact that most of the genes in any cell are expected to be irrelevant to the specific context in hand, as well as the fact that many genes might participate in multiple processes. This has been achieved by clustering the given set of genes while allowing any gene to have any of the three eventualities, to be exclusively assigned to a single cluster, to be simultaneously assigned to multiple clusters, or not to be assigned to any of the clusters. In this study, we expand the scope of application of the Bi-CoPaM method by applying it, for the first time, to bacterial datasets, namely to a set of five Escherichia coli bacterial datasets generated under different biological conditions, in order to identify the subsets of genes which are consistently co-expressed, i.e. well correlated with each other. We identify two clusters with such consistent co-expression, and interestingly, they themselves are consistently negatively correlated with each other. The first cluster is enriched with genes participating in protein synthesis and DNA repair while the second is enriched with transporting genes. Consequently, we draw biological hypotheses that relate some of the genes with currently unknown biological processes to their potential processes. These hypotheses can serve as pilots for focused future gene discovery studies.
-
MLSP - Method for the identification of the subsets of genes specifically consistently co-expressed in a set of datasets
2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), 2013Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The recently proposed binarization of Consensus Partition matrices (Bi-CoPaM) ensemble clustering method has offered the ability to mine multiple genome-wide microarray datasets for the subsets of genes which are consistently co-expressed in all of these datasets. Though, some of those subsets of genes might also be consistently co-expressed in many other datasets that were generated under a wider range of conditions than those of interest in a single focused study. Here we propose a new method, named as the unification of clustering results from multiple datasets using external specifications (UNCLES). The external specifications imposed in this study aim at mining for the subsets of genes that are consistently co-expressed in one set of datasets (S+) and not consistently co-expressed in another set of datasets (S-). We tested our proposed method over eight budding yeast cell-cycle datasets for S+ and other six general budding yeast datasets for S-. Our results have shown the ability of our method to find the subsets of genes consistently co-expressed in the S+ datasets successfully, while excluding the subsets of genes that are also consistently co-expressed in the S- datasets.
David J Roberts - One of the best experts on this subject based on the ideXlab platform.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets. Results Here, we propose an unsupervised method for the unification of clustering results from multiple datasets using external specifications (UNCLES) . This method has the ability to identify the subsets of genes consistently co-expressed in a subset of datasets while being poorly co-expressed in another subset of datasets, and to identify the subsets of genes consistently co-expressed in all given datasets. We also propose the M-N scatter plots validation technique and adopt it to set the parameters of UNCLES, such as the number of clusters, automatically. Additionally, we propose an approach for the synthesis of gene expression datasets using real data profiles in a way which combines the ground-truth-knowledge of synthetic data and the realistic expression values of real data, and therefore overcomes the problem of faithfulness of synthetic expression data modelling. By application to those datasets, we validate UNCLES while comparing it with other conventional clustering methods, and of particular relevance, biclustering methods. We further validate UNCLES by application to a set of 14 real genome-wide yeast datasets as it produces focused clusters that conform well to known biological facts. Furthermore, in-silico -based hypotheses regarding the function of a few previously unknown genes in those focused clusters are drawn. Conclusions The UNCLES method, the M-N scatter plots technique, and the expression data synthesis approach will have wide application for the comprehensive analysis of genomic and other sources of multiple complex biological datasets. Moreover, the derived in-silico -based biological hypotheses represent subjects for future functional studies.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets.
-
Application of the Bi-CoPaM Method to Five Escherichia Coli Datasets Generated under Various Biological Conditions
Journal of Signal Processing Systems, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The increasing amounts of high-throughput biological datasets stimulate the information engineering and machine learning research community to direct more studies towards designing and applying novel methods which are sophisticated and specialised to tackle the problems that are specific in such datasets. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method tackles the problem of scrutinising multiple gene expression microarray datasets to identify the subsets of genes which are consistently co-expressed across them. It allows for clustering results which better reflect the biological fact that most of the genes in any cell are expected to be irrelevant to the specific context in hand, as well as the fact that many genes might participate in multiple processes. This has been achieved by clustering the given set of genes while allowing any gene to have any of the three eventualities, to be exclusively assigned to a single cluster, to be simultaneously assigned to multiple clusters, or not to be assigned to any of the clusters. In this study, we expand the scope of application of the Bi-CoPaM method by applying it, for the first time, to bacterial datasets, namely to a set of five Escherichia coli bacterial datasets generated under different biological conditions, in order to identify the subsets of genes which are consistently co-expressed, i.e. well correlated with each other. We identify two clusters with such consistent co-expression, and interestingly, they themselves are consistently negatively correlated with each other. The first cluster is enriched with genes participating in protein synthesis and DNA repair while the second is enriched with transporting genes. Consequently, we draw biological hypotheses that relate some of the genes with currently unknown biological processes to their potential processes. These hypotheses can serve as pilots for focused future gene discovery studies.
-
MLSP - Method for the identification of the subsets of genes specifically consistently co-expressed in a set of datasets
2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), 2013Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The recently proposed binarization of Consensus Partition matrices (Bi-CoPaM) ensemble clustering method has offered the ability to mine multiple genome-wide microarray datasets for the subsets of genes which are consistently co-expressed in all of these datasets. Though, some of those subsets of genes might also be consistently co-expressed in many other datasets that were generated under a wider range of conditions than those of interest in a single focused study. Here we propose a new method, named as the unification of clustering results from multiple datasets using external specifications (UNCLES). The external specifications imposed in this study aim at mining for the subsets of genes that are consistently co-expressed in one set of datasets (S+) and not consistently co-expressed in another set of datasets (S-). We tested our proposed method over eight budding yeast cell-cycle datasets for S+ and other six general budding yeast datasets for S-. Our results have shown the ability of our method to find the subsets of genes consistently co-expressed in the S+ datasets successfully, while excluding the subsets of genes that are also consistently co-expressed in the S- datasets.
-
ICASSP - Identification of genes consistently co-expressed in multiple microarray datasets by a genome-wide Bi-CoPaM approach
2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Many methods have been proposed to identify informative subsets of genes in microarray studies in order to focus the research. For instance, the recently proposed binarization of Consensus Partition matrices (Bi-CoPaM) method has, amongst its various features, the ability to generate tight clusters of genes while leaving many genes unassigned from all clusters. We propose exploiting this particular feature by applying the Bi-CoPaM over genome-wide microarray data from multiple datasets to generate more clusters than required. Then, these clusters are tightened so that most of their genes are left unassigned from all clusters, and most of the clusters are left totally empty. The tightened clusters, which are still not empty, include those genes that are consistently co-expressed in multiple datasets when examined by various clustering methods. An example of this is demonstrated in this paper for cyclic and acyclic genes as well as for genes that are highly expressed and that are not. Thus, the results of our proposed approach cannot be reproduced by other methods of genes' periodicity identification or by other methods of clustering.
Marcello Pelillo - One of the best experts on this subject based on the ideXlab platform.
-
Probabilistic Consensus clustering using evidence accumulation
Machine Learning, 2015Co-Authors: André Lourenço, N. Rebagliati, Samuel Rota Bulò, Ana L. N. Fred, Mario A T Figueiredo, Marcello PelilloAbstract:Clustering ensemble methods produce a Consensus Partition of a set of data points by combining the results of a collection of base clustering algorithms. In the evidence accumulation clustering (EAC) paradigm, the clustering ensemble is transformed into a pairwise co-association matrix, thus avoiding the label correspondence problem, which is intrinsic to other clustering ensemble schemes. In this paper, we propose a Consensus clustering approach based on the EAC paradigm, which is not limited to crisp Partitions and fully exploits the nature of the co-association matrix. Our solution determines probabilistic assignments of data points to clusters by minimizing a Bregman divergence between the observed co-association frequencies and the corresponding co-occurrence probabilities expressed as functions of the unknown assignments. We additionally propose an optimization algorithm to find a solution under any double-convex Bregman divergence. Experiments on both synthetic and real benchmark data show the effectiveness of the proposed approach.
-
ICPRAM (Selected Papers) - A MAP approach to evidence accumulation clustering
Advances in Intelligent Systems and Computing, 2014Co-Authors: André Lourenço, N. Rebagliati, Mario A T Figueiredo, Ana Fred, Samuel Rota Bulò, Marcello PelilloAbstract:The Evidence Accumulation Clustering (EAC) paradigm is a clustering ensemble method which derives a Consensus Partition from a collection of base clusterings obtained using different algorithms. It collects from the Partitions in the ensemble a set of pairwise observations about the co-occurrence of objects in a same cluster and it uses these co-occurrence statistics to derive a similarity matrix, referred to as co-association matrix. The Probabilistic Evidence Accumulation for Clustering Ensembles (PEACE) algorithm is a principled approach for the extraction of a Consensus clustering from the observations encoded in the co-association matrix based on a probabilistic model for the co-association matrix parameterized by the unknown assignments of objects to clusters. In this paper we extend the PEACE algorithm by deriving a Consensus solution according to a MAP approach with Dirichlet priors defined for the unknown probabilistic cluster assignments. In particular, we study the positive regularization effect of Dirichlet priors on the final Consensus solution with both synthetic and real benchmark data.
-
EMMCVPR - Consensus Clustering with Robust Evidence Accumulation
Lecture Notes in Computer Science, 2013Co-Authors: André Lourenço, Ana Fred, Samuel Rota Bulò, Marcello PelilloAbstract:Consensus clustering methodologies combine a set of Partitions on the clustering ensemble providing a Consensus Partition. One of the drawbacks of the standard combination algorithms is that all the Partitions of the ensemble have the same weight on the aggregation process. By making a differentiation among the Partitions the quality of the Consensus could be improved. In this paper we propose a novel formulation that tries to find a median-Partition for the clustering ensemble process based on the evidence accumulation framework, but including a weighting mechanism that allows to differentiate the importance of the Partitions of the ensemble in order to become more robust to noisy ensembles. Experiments on both synthetic and real benchmark data show the effectiveness of the proposed approach.
-
Probabilistic Consensus clustering using evidence accumulation
Machine Learning, 2013Co-Authors: André Lourenço, N. Rebagliati, Samuel Rota Bulò, Ana L. N. Fred, Mario A T Figueiredo, Marcello PelilloAbstract:© 2013, The Author(s). Clustering ensemble methods produce a Consensus Partition of a set of data points by combining the results of a collection of base clustering algorithms. In the evidence accumulation clustering (EAC) paradigm, the clustering ensemble is transformed into a pairwise co-association matrix, thus avoiding the label correspondence problem, which is intrinsic to other clustering ensemble schemes. In this paper, we propose a Consensus clustering approach based on the EAC paradigm, which is not limited to crisp Partitions and fully exploits the nature of the co-association matrix. Our solution determines probabilistic assignments of data points to clusters by minimizing a Bregman divergence between the observed co-association frequencies and the corresponding co-occurrence probabilities expressed as functions of the unknown assignments. We additionally propose an optimization algorithm to find a solution under any double-convex Bregman divergence. Experiments on both synthetic and real benchmark data show the effectiveness of the proposed approach.
Rui Fa - One of the best experts on this subject based on the ideXlab platform.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets. Results Here, we propose an unsupervised method for the unification of clustering results from multiple datasets using external specifications (UNCLES) . This method has the ability to identify the subsets of genes consistently co-expressed in a subset of datasets while being poorly co-expressed in another subset of datasets, and to identify the subsets of genes consistently co-expressed in all given datasets. We also propose the M-N scatter plots validation technique and adopt it to set the parameters of UNCLES, such as the number of clusters, automatically. Additionally, we propose an approach for the synthesis of gene expression datasets using real data profiles in a way which combines the ground-truth-knowledge of synthetic data and the realistic expression values of real data, and therefore overcomes the problem of faithfulness of synthetic expression data modelling. By application to those datasets, we validate UNCLES while comparing it with other conventional clustering methods, and of particular relevance, biclustering methods. We further validate UNCLES by application to a set of 14 real genome-wide yeast datasets as it produces focused clusters that conform well to known biological facts. Furthermore, in-silico -based hypotheses regarding the function of a few previously unknown genes in those focused clusters are drawn. Conclusions The UNCLES method, the M-N scatter plots technique, and the expression data synthesis approach will have wide application for the comprehensive analysis of genomic and other sources of multiple complex biological datasets. Moreover, the derived in-silico -based biological hypotheses represent subjects for future functional studies.
-
UNCLES: method for the identification of genes differentially consistently co-expressed in a specific subset of datasets
BMC Bioinformatics, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Background Collective analysis of the increasingly emerging gene expression datasets are required. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method can combine clustering results from multiple datasets to identify the subsets of genes which are consistently co-expressed in all of the provided datasets in a tuneable manner. However, results validation and parameter setting are issues that complicate the design of such methods. Moreover, although it is a common practice to test methods by application to synthetic datasets, the mathematical models used to synthesise such datasets are usually based on approximations which may not always be sufficiently representative of real datasets.
-
Application of the Bi-CoPaM Method to Five Escherichia Coli Datasets Generated under Various Biological Conditions
Journal of Signal Processing Systems, 2015Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The increasing amounts of high-throughput biological datasets stimulate the information engineering and machine learning research community to direct more studies towards designing and applying novel methods which are sophisticated and specialised to tackle the problems that are specific in such datasets. The recently proposed binarisation of Consensus Partition matrices (Bi-CoPaM) method tackles the problem of scrutinising multiple gene expression microarray datasets to identify the subsets of genes which are consistently co-expressed across them. It allows for clustering results which better reflect the biological fact that most of the genes in any cell are expected to be irrelevant to the specific context in hand, as well as the fact that many genes might participate in multiple processes. This has been achieved by clustering the given set of genes while allowing any gene to have any of the three eventualities, to be exclusively assigned to a single cluster, to be simultaneously assigned to multiple clusters, or not to be assigned to any of the clusters. In this study, we expand the scope of application of the Bi-CoPaM method by applying it, for the first time, to bacterial datasets, namely to a set of five Escherichia coli bacterial datasets generated under different biological conditions, in order to identify the subsets of genes which are consistently co-expressed, i.e. well correlated with each other. We identify two clusters with such consistent co-expression, and interestingly, they themselves are consistently negatively correlated with each other. The first cluster is enriched with genes participating in protein synthesis and DNA repair while the second is enriched with transporting genes. Consequently, we draw biological hypotheses that relate some of the genes with currently unknown biological processes to their potential processes. These hypotheses can serve as pilots for focused future gene discovery studies.
-
MLSP - Method for the identification of the subsets of genes specifically consistently co-expressed in a set of datasets
2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), 2013Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:The recently proposed binarization of Consensus Partition matrices (Bi-CoPaM) ensemble clustering method has offered the ability to mine multiple genome-wide microarray datasets for the subsets of genes which are consistently co-expressed in all of these datasets. Though, some of those subsets of genes might also be consistently co-expressed in many other datasets that were generated under a wider range of conditions than those of interest in a single focused study. Here we propose a new method, named as the unification of clustering results from multiple datasets using external specifications (UNCLES). The external specifications imposed in this study aim at mining for the subsets of genes that are consistently co-expressed in one set of datasets (S+) and not consistently co-expressed in another set of datasets (S-). We tested our proposed method over eight budding yeast cell-cycle datasets for S+ and other six general budding yeast datasets for S-. Our results have shown the ability of our method to find the subsets of genes consistently co-expressed in the S+ datasets successfully, while excluding the subsets of genes that are also consistently co-expressed in the S- datasets.
-
ICASSP - Identification of genes consistently co-expressed in multiple microarray datasets by a genome-wide Bi-CoPaM approach
2013 IEEE International Conference on Acoustics Speech and Signal Processing, 2013Co-Authors: Basel Abu-jamous, Rui Fa, David J Roberts, Asoke K NandiAbstract:Many methods have been proposed to identify informative subsets of genes in microarray studies in order to focus the research. For instance, the recently proposed binarization of Consensus Partition matrices (Bi-CoPaM) method has, amongst its various features, the ability to generate tight clusters of genes while leaving many genes unassigned from all clusters. We propose exploiting this particular feature by applying the Bi-CoPaM over genome-wide microarray data from multiple datasets to generate more clusters than required. Then, these clusters are tightened so that most of their genes are left unassigned from all clusters, and most of the clusters are left totally empty. The tightened clusters, which are still not empty, include those genes that are consistently co-expressed in multiple datasets when examined by various clustering methods. An example of this is demonstrated in this paper for cyclic and acyclic genes as well as for genes that are highly expressed and that are not. Thus, the results of our proposed approach cannot be reproduced by other methods of genes' periodicity identification or by other methods of clustering.
Yun Fu - One of the best experts on this subject based on the ideXlab platform.
-
IJCAI - Adversarial Graph Embedding for Ensemble Clustering
Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, 2019Co-Authors: Jun Li, Zhaowen Wang, Yun FuAbstract:Ensemble clustering generally integrates basic Partitions into a Consensus one through a graph Partitioning method, which, however, has two limitations: 1) it neglects to reuse original features; 2) obtaining Consensus Partition with learnable graph representations is still under-explored. In this paper, we propose a novel Adversarial Graph Auto-Encoders (AGAE) model to incorporate ensemble clustering into a deep graph embedding process. Specifically, graph convolutional network is adopted as probabilistic encoder to jointly integrate the information from feature content and Consensus graph, and a simple inner product layer is used as decoder to reconstruct graph with the encoded latent variables (i.e., embedding representations). Moreover, we develop an adversarial regularizer to guide the network training with an adaptive Partition-dependent prior. Experiments on eight real-world datasets are presented to show the effectiveness of AGAE over several state-of-the-art deep embedding and ensemble clustering methods.
-
Robust Spectral Ensemble Clustering via Rank Minimization
ACM Transactions on Knowledge Discovery From Data, 2019Co-Authors: Sheng Li, Zhengming Ding, Yun FuAbstract:Ensemble Clustering (EC) is an important topic for data cluster analysis. It targets to integrate multiple Basic Partitions (BPs) of a particular dataset into a Consensus Partition. Among previous works, one promising and effective way is to transform EC as a graph Partitioning problem on the co-association matrix, which is a pair-wise similarity matrix summarized by all the BPs in essence. However, most existing EC methods directly utilize the co-association matrix, yet without considering various noises (e.g., the disagreement between different BPs and the outliers) that may exist in it. These noises can impair the cluster structure of a co-association matrix, and thus mislead the final graph Partitioning process. To address this challenge, we propose a novel Robust Spectral Ensemble Clustering (RSEC) algorithm in this article. Specifically, we learn low-rank representation (LRR) for the co-association matrix to uncover its cluster structure and handle the noises, and meanwhile, we perform spectral clustering with the learned representation to seek for a Consensus Partition. These two steps are jointly proceeded within a unified optimization framework. In particular, during the optimizing process, we leverage Consensus Partition to iteratively enhance the block-diagonal structure of LRR, in order to assist the graph Partitioning. To solve RSEC, we first formulate it by using nuclear norm as a convex proxy to the rank function. Then, motivated by the recent advances in non-convex rank minimization, we further develop a non-convex model for RSEC and provide it a solution by the majorization--minimization Augmented Lagrange Multiplier algorithm. Experiments on 18 real-world datasets demonstrate the effectiveness of our algorithm compared with state-of-the-art methods. Moreover, several impact factors on the clustering performance of our approach are also explored extensively.
-
IJCAI - From Ensemble Clustering to Multi-View Clustering.
Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 2017Co-Authors: Sheng Li, Zhengming Ding, Yun FuAbstract:Multi-View Clustering (MVC) aims to find the cluster structure shared by multiple views of a particular dataset. Existing MVC methods mainly integrate the raw data from different views, while ignoring the high-level information. Thus, their performance may degrade due to the conflict between heterogeneous features and the noises existing in each individual view. To overcome this problem, we propose a novel Multi-View Ensemble Clustering (MVEC) framework to solve MVC in an Ensemble Clustering (EC) way, which generates Basic Partitions (BPs) for each view individually and seeks for a Consensus Partition among all the BPs. By this means, we naturally leverage the complementary information of multi-view data in the same Partition space. Instead of directly fusing BPs, we employ the low-rank and sparse decomposition to explicitly consider the connection between different views and detect the noises in each view. Moreover, the spectral ensemble clustering task is also involved by our framework with a carefully designed constraint, making MVEC a unified optimization framework to achieve the final Consensus Partition. Experimental results on six real-world datasets show the efficacy of our approach compared with both MVC and EC methods.
-
From ensemble clustering to multi-view clustering
IJCAI International Joint Conference on Artificial Intelligence, 2017Co-Authors: Zhiqiang Tao, Hongfu Liu, Zhengming Ding, Sheng Li, Yun FuAbstract:Multi-View Clustering (MVC) aims to find the cluster structure shared by multiple views of a par-ticular dataset. Existing MVC methods mainly in-tegrate the raw data from different views, while ignoring the high-level information. Thus, their performance may degrade due to the conflict be-tween heterogeneous features and the noises ex-isting in each individual view. To overcome this problem, we propose a novel Multi-View Ensem-ble Clustering (MVEC) framework to solve MVC in an Ensemble Clustering (EC) way, which gener-ates Basic Partitions (BPs) for each view individu-ally and seeks for a Consensus Partition among all the BPs. By this means, we naturally leverage the complementary information of multi-view data in the same Partition space. Instead of directly fusing BPs, we employ the low-rank and sparse decom-position to explicitly consider the connection be-tween different views and detect the noises in each view. Moreover, the spectral ensemble clustering task is also involved by our framework with a care-fully designed constraint, making MVEC a unified optimization framework to achieve the final con-sensus Partition. Experimental results on six real-world datasets show the efficacy of our approach compared with both MVC and EC methods.
-
CIKM - Robust Spectral Ensemble Clustering
Proceedings of the 25th ACM International on Conference on Information and Knowledge Management, 2016Co-Authors: Sheng Li, Yun FuAbstract:Ensemble Clustering (EC) aims to integrate multiple Basic Partitions (BPs) of the same dataset into a Consensus one. It could be transformed as a graph Partition problem on the co-association matrix derived from BPs. However, existing EC methods usually directly use the co-association matrix, yet without considering various noises (e.g., the disagreement between different BPs or outliers) that may exist in it. These noises can impair the cluster structure of a co-association matrix and thus degrade the final clustering performance. In this paper, we propose a novel Robust Spectral Ensemble Clustering (RSEC) approach to address this challenge. First, RSEC learns a robust representation for the co-association matrix through low-rank constraint, which reveals the cluster structure of a co-association matrix and captures various noises in it. Second, RSEC finds the Consensus Partition by conducting spectral clustering. These two steps are iteratively performed in a unified optimization framework. Most importantly, during our optimization process, we utilize Consensus Partition to iteratively enhance the block-diagonal structure of the learned representation to further assist the clustering process. Experiments on numerous real-world datasets demonstrate the effectiveness of our method compared with the state-of-the-art. Moreover, several impact factors that may affect the clustering performance of our approach are also explored extensively.