The Experts below are selected from a list of 109812 Experts worldwide ranked by ideXlab platform

Vipin Kumar - One of the best experts on this subject based on the ideXlab platform.

  • Discovery of error-tolerant biclusters from noisy Gene Expression Data
    BMC Bioinformatics, 2011
    Co-Authors: Rohit Gupta, Vipin Kumar
    Abstract:

    Background An important analysis performed on microarray Gene-Expression Data is to discover biclusters, which denote groups of Genes that are coherently expressed for a subset of conditions. Various biclustering algorithms have been proposed to find different types of biclusters from these real-valued Gene-Expression Data sets. However, these algorithms suffer from several limitations such as inability to explicitly handle errors/noise in the Data; difficulty in discovering small bicliusters due to their top-down approach; inability of some of the approaches to find overlapping biclusters, which is crucial as many Genes participate in multiple biological processes. Association pattern mining also produce biclusters as their result and can naturally address some of these limitations. However, traditional association mining only finds exact biclusters, which limits its applicability in real-life Data sets where the biclusters may be fragmented due to random noise/errors. Moreover, as they only work with binary or boolean attributes, their application on Gene-Expression Data require transforming real-valued attributes to binary attributes, which often results in loss of information. Many past approaches have tried to address the issue of noise and handling real-valued attributes independently but there is no systematic approach that addresses both of these issues together. Results In this paper, we first propose a novel error-tolerant biclustering model, ‘ ET-bicluster ’, and then propose a bottom-up heuristic-based mining algorithm to sequentially discover error-tolerant biclusters directly from real-valued Gene-Expression Data. The efficacy of our proposed approach is illustrated by comparing it with a recent approach RAP in the context of two biological problems: discovery of functional modules and discovery of biomarkers. For the first problem, two real-valued S.Cerevisiae microarray Gene-Expression Data sets are used to demonstrate that the biclusters obtained from ET-bicluster approach not only recover larger set of Genes as compared to those obtained from RAP approach but also have higher functional coherence as evaluated using the GO-based functional enrichment analysis. The statistical significance of the discovered error-tolerant biclusters as estimated by using two randomization tests, reveal that they are indeed biologically meaningful and statistically significant. For the second problem of biomarker discovery, we used four real-valued Breast Cancer microarray Gene-Expression Data sets and evaluate the biomarkers obtained using MSigDB Gene sets. Conclusions The results obtained for both the problems: functional module discovery and biomarkers discovery, clearly signifies the usefulness of the proposed ET-bicluster approach and illustrate the importance of explicitly incorporating noise/errors in discovering coherent groups of Genes from Gene-Expression Data.

  • Discovery of error-tolerant biclusters from noisy Gene Expression Data.
    BMC bioinformatics, 2011
    Co-Authors: Rohit Gupta, Navneet Rao, Vipin Kumar
    Abstract:

    BACKGROUND: An important analysis performed on microarray Gene-Expression Data is to discover biclusters, which denote groups of Genes that are coherently expressed for a subset of conditions. Various biclustering algorithms have been proposed to find different types of biclusters from these real-valued Gene-Expression Data sets. However, these algorithms suffer from several limitations such as inability to explicitly handle errors/noise in the Data; difficulty in discovering small bicliusters due to their top-down approach; inability of some of the approaches to find overlapping biclusters, which is crucial as many Genes participate in multiple biological processes. Association pattern mining also produce biclusters as their result and can naturally address some of these limitations. However, traditional association mining only finds exact biclusters, which limits its applicability in real-life Data sets where the biclusters may be fragmented due to random noise/errors. Moreover, as they only work with binary or boolean attributes, their application on Gene-Expression Data require transforming real-valued attributes to binary attributes, which often results in loss of information. Many past approaches have tried to address the issue of noise and handling real-valued attributes independently but there is no systematic approach that addresses both of these issues together.\n\nRESULTS: In this paper, we first propose a novel error-tolerant biclustering model, 'ET-bicluster', and then propose a bottom-up heuristic-based mining algorithm to sequentially discover error-tolerant biclusters directly from real-valued Gene-Expression Data. The efficacy of our proposed approach is illustrated by comparing it with a recent approach RAP in the context of two biological problems: discovery of functional modules and discovery of biomarkers. For the first problem, two real-valued S.Cerevisiae microarray Gene-Expression Data sets are used to demonstrate that the biclusters obtained from ET-bicluster approach not only recover larger set of Genes as compared to those obtained from RAP approach but also have higher functional coherence as evaluated using the GO-based functional enrichment analysis. The statistical significance of the discovered error-tolerant biclusters as estimated by using two randomization tests, reveal that they are indeed biologically meaningful and statistically significant. For the second problem of biomarker discovery, we used four real-valued Breast Cancer microarray Gene-Expression Data sets and evaluate the biomarkers obtained using MSigDB Gene sets.\n\nCONCLUSIONS: The results obtained for both the problems: functional module discovery and biomarkers discovery, clearly signifies the usefulness of the proposed ET-bicluster approach and illustrate the importance of explicitly incorporating noise/errors in discovering coherent groups of Genes from Gene-Expression Data.

  • BIBM - Systematic Evaluation of Scaling Methods for Gene Expression Data
    2008 IEEE International Conference on Bioinformatics and Biomedicine, 2008
    Co-Authors: Gaurav Pandey, Lakshmi Naarayanan Ramakrishnan, Michael Steinbach, Vipin Kumar
    Abstract:

    Even after an experimentally prepared Gene Expression Data set has been pre-processed to account for variations in the microarray technology, there may be inconsistencies between the scales of measurements in different conditions. This may happen for reasons such as the accumulation of Gene Expression Data prepared by different laboratories into a single Data set. A variety of scaling and transformation methods have been used for addressing these scale inconsistencies in different studies on the analysis of Gene Expression Data sets. However, a quantitative estimation of their relative performance has been lacking. In this paper, we report an extensive evaluation of scaling and transformation methods for their effectiveness with respect to the important problem of protein function prediction. We consider several such commonly used methods for Gene Expression Data, such as z-score scaling, quantile normalization, diff transformation, and two new scaling methods, sigmoid and double sigmoid, that have not been used previously in this domain to the best of our knowledge. We show that the performance of these methods can vary significantly across Data sets, but Dsigmoid scaling and z-score transformation Generally perform well for the two types of Gene Expression Data, namely temporal and non-temporal, respectively.

  • Systematic Evaluation of Scaling Methods for Gene Expression Data
    2008 IEEE International Conference on Bioinformatics and Biomedicine, 2008
    Co-Authors: Gaurav Pandey, Lakshmi Naarayanan Ramakrishnan, Michael Steinbach, Vipin Kumar
    Abstract:

    Even after an experimentally prepared Gene Expression Data set has been pre-processed to account for variations in the microarray technology, there may be inconsistencies between the scales of measurements in different conditions. This may happen for reasons such as the accumulation of Gene Expression Data prepared by different laboratories into a single Data set. A variety of scaling and transformation methods have been used for addressing these scale inconsistencies in different studies on the analysis of Gene Expression Data sets. However, a quantitative estimation of their relative performance has been lacking. In this paper, we report an extensive evaluation of scaling and transformation methods for their effectiveness with respect to the important problem of protein function prediction. We consider several such commonly used methods for Gene Expression Data, such as z-score scaling, quantile normalization, diff transformation, and two new scaling methods, sigmoid and double sigmoid, that have not been used previously in this domain to the best of our knowledge. We show that the performance of these methods can vary significantly across Data sets, but Dsigmoid scaling and z-score transformation Generally perform well for the two types of Gene Expression Data, namely temporal and non-temporal, respectively.

Rohit Gupta - One of the best experts on this subject based on the ideXlab platform.

  • Discovery of error-tolerant biclusters from noisy Gene Expression Data
    BMC Bioinformatics, 2011
    Co-Authors: Rohit Gupta, Vipin Kumar
    Abstract:

    Background An important analysis performed on microarray Gene-Expression Data is to discover biclusters, which denote groups of Genes that are coherently expressed for a subset of conditions. Various biclustering algorithms have been proposed to find different types of biclusters from these real-valued Gene-Expression Data sets. However, these algorithms suffer from several limitations such as inability to explicitly handle errors/noise in the Data; difficulty in discovering small bicliusters due to their top-down approach; inability of some of the approaches to find overlapping biclusters, which is crucial as many Genes participate in multiple biological processes. Association pattern mining also produce biclusters as their result and can naturally address some of these limitations. However, traditional association mining only finds exact biclusters, which limits its applicability in real-life Data sets where the biclusters may be fragmented due to random noise/errors. Moreover, as they only work with binary or boolean attributes, their application on Gene-Expression Data require transforming real-valued attributes to binary attributes, which often results in loss of information. Many past approaches have tried to address the issue of noise and handling real-valued attributes independently but there is no systematic approach that addresses both of these issues together. Results In this paper, we first propose a novel error-tolerant biclustering model, ‘ ET-bicluster ’, and then propose a bottom-up heuristic-based mining algorithm to sequentially discover error-tolerant biclusters directly from real-valued Gene-Expression Data. The efficacy of our proposed approach is illustrated by comparing it with a recent approach RAP in the context of two biological problems: discovery of functional modules and discovery of biomarkers. For the first problem, two real-valued S.Cerevisiae microarray Gene-Expression Data sets are used to demonstrate that the biclusters obtained from ET-bicluster approach not only recover larger set of Genes as compared to those obtained from RAP approach but also have higher functional coherence as evaluated using the GO-based functional enrichment analysis. The statistical significance of the discovered error-tolerant biclusters as estimated by using two randomization tests, reveal that they are indeed biologically meaningful and statistically significant. For the second problem of biomarker discovery, we used four real-valued Breast Cancer microarray Gene-Expression Data sets and evaluate the biomarkers obtained using MSigDB Gene sets. Conclusions The results obtained for both the problems: functional module discovery and biomarkers discovery, clearly signifies the usefulness of the proposed ET-bicluster approach and illustrate the importance of explicitly incorporating noise/errors in discovering coherent groups of Genes from Gene-Expression Data.

  • Discovery of error-tolerant biclusters from noisy Gene Expression Data.
    BMC bioinformatics, 2011
    Co-Authors: Rohit Gupta, Navneet Rao, Vipin Kumar
    Abstract:

    BACKGROUND: An important analysis performed on microarray Gene-Expression Data is to discover biclusters, which denote groups of Genes that are coherently expressed for a subset of conditions. Various biclustering algorithms have been proposed to find different types of biclusters from these real-valued Gene-Expression Data sets. However, these algorithms suffer from several limitations such as inability to explicitly handle errors/noise in the Data; difficulty in discovering small bicliusters due to their top-down approach; inability of some of the approaches to find overlapping biclusters, which is crucial as many Genes participate in multiple biological processes. Association pattern mining also produce biclusters as their result and can naturally address some of these limitations. However, traditional association mining only finds exact biclusters, which limits its applicability in real-life Data sets where the biclusters may be fragmented due to random noise/errors. Moreover, as they only work with binary or boolean attributes, their application on Gene-Expression Data require transforming real-valued attributes to binary attributes, which often results in loss of information. Many past approaches have tried to address the issue of noise and handling real-valued attributes independently but there is no systematic approach that addresses both of these issues together.\n\nRESULTS: In this paper, we first propose a novel error-tolerant biclustering model, 'ET-bicluster', and then propose a bottom-up heuristic-based mining algorithm to sequentially discover error-tolerant biclusters directly from real-valued Gene-Expression Data. The efficacy of our proposed approach is illustrated by comparing it with a recent approach RAP in the context of two biological problems: discovery of functional modules and discovery of biomarkers. For the first problem, two real-valued S.Cerevisiae microarray Gene-Expression Data sets are used to demonstrate that the biclusters obtained from ET-bicluster approach not only recover larger set of Genes as compared to those obtained from RAP approach but also have higher functional coherence as evaluated using the GO-based functional enrichment analysis. The statistical significance of the discovered error-tolerant biclusters as estimated by using two randomization tests, reveal that they are indeed biologically meaningful and statistically significant. For the second problem of biomarker discovery, we used four real-valued Breast Cancer microarray Gene-Expression Data sets and evaluate the biomarkers obtained using MSigDB Gene sets.\n\nCONCLUSIONS: The results obtained for both the problems: functional module discovery and biomarkers discovery, clearly signifies the usefulness of the proposed ET-bicluster approach and illustrate the importance of explicitly incorporating noise/errors in discovering coherent groups of Genes from Gene-Expression Data.

Aidong Zhang - One of the best experts on this subject based on the ideXlab platform.

  • An interactive approach to mining Gene Expression Data
    IEEE Transactions on Knowledge and Data Engineering, 2005
    Co-Authors: Daxin Jiang, Aidong Zhang
    Abstract:

    Effective identification of coexpressed Genes and coherent patterns in Gene Expression Data is an important task in bioinformatics research and biomedical applications. Several clustering methods have recently been proposed to identify coexpressed Genes that share similar coherent patterns. However, there is no objective standard for groups of coexpressed Genes. The interpretation of co-Expression heavily depends on domain knowledge. Furthermore, groups of coexpressed Genes in Gene Expression Data are often highly connected through a large number of "intermediate" Genes. There may be no clear boundaries to separate clusters. Clustering Gene Expression Data also faces the challenges of satisfying biological domain requirements and addressing the high connectivity of the Data sets. In this paper, we propose an interactive framework for exploring coherent patterns in Gene Expression Data. A novel coherent pattern index is proposed to give users highly confident indications of the existence of coherent patterns. To derive a coherent pattern index and facilitate clustering, we devise an attraction tree structure that summarizes the coherence information among Genes in the Data set. We present efficient and scalable algorithms for constructing attraction trees and coherent pattern indices from Gene Expression Data sets. Our experimental results show that our approach is effective in mining Gene Expression Data and is scalable for mining large Data sets.

  • VLDB - GPX: interactive mining of Gene Expression Data
    Proceedings 2004 VLDB Conference, 2004
    Co-Authors: Daxin Jiang, Aidong Zhang
    Abstract:

    Discovering co-expressed Genes and coherent Expression patterns in Gene Expression Data is an important Data analysis task in bioinformatics research and biomedical applications. Although various clustering methods have been proposed, two tough challenges still remain on how to integrate the users' domain knowledge and how to handle the high connectivity in the Data. Recently, we have systematically studied the problem and proposed an effective approach [3]. In this paper, we describe a demonstration of GPX (for Gene Pattern eXplorer), an integrated environment for interactive exploration of coherent Expression patterns and co-expressed Genes in Gene Expression Data. GPX integrates several novel techniques, including the coherent pattern index graph, a Gene annotation panel, and a graphical interface, to adopt users' domain knowledge and support explorative operations in the clustering procedure. The GPX system as well as its techniques will be showcased, and the progress of GPX will be exemplified using several real-world Gene Expression Data sets.

  • Cluster analysis for Gene Expression Data: A survey
    IEEE Transactions on Knowledge and Data Engineering, 2004
    Co-Authors: Daxin Jiang, Chun Tang, Aidong Zhang
    Abstract:

    DNA microarray technology has now made it possible to simultaneously monitor the Expression levels of thousands of Genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in Gene Expression Data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of Genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of Data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the Data mining process to reveal natural structures and identify interesting patterns in the underlying Data. Cluster analysis seeks to partition a given Data set into groups based on specified features so that the Data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to Gene Expression Data, and also new algorithms have recently been proposed specifically aiming at Gene Expression Data. These clustering algorithms have been proven useful for identifying biologically relevant groups of Genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on Gene Expression Data. In particular, we divide cluster analysis for Gene Expression Data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.

Boyun Zhang - One of the best experts on this subject based on the ideXlab platform.

  • SVM-Based Tumor Classification with Gene Expression Data
    Lecture Notes in Computer Science, 2020
    Co-Authors: Shulin Wang, Ji Wang, Huowang Chen, Boyun Zhang
    Abstract:

    Gene Expression Data that are gathered from tissue samples are expected to significantly help the development of efficient tumor diagnosis and classification platforms. Since DNA microarray experiments provide us with huge amount of Gene Expression Data and only a few of Genes are related to tumor, Gene selection algorithms should be emphatically explored to extract those informative Genes related tumor from Gene Expression Data. So we propose a novel feature selection approach to further improve the SVM-based classification performance of Gene Expression Data, which projects high dimensional Data onto lower dimensional feature space. We examine a set of Gene Expression Data that include sets of tumor and normal clinical samples by means of SVMs classifier. Experiments show that SVM has a superior performance in classification of Gene Expression Data as long as the selected features can represent the principal components of all Gene Expression samples.

  • ADMA - SVM-Based tumor classification with Gene Expression Data
    Advanced Data Mining and Applications, 2006
    Co-Authors: Shulin Wang, Ji Wang, Huowang Chen, Boyun Zhang
    Abstract:

    Gene Expression Data that are gathered from tissue samples are expected to significantly help the development of efficient tumor diagnosis and classification platforms. Since DNA microarray experiments provide us with huge amount of Gene Expression Data and only a few of Genes are related to tumor, Gene selection algorithms should be emphatically explored to extract those informative Genes related tumor from Gene Expression Data. So we propose a novel feature selection approach to further improve the SVM-based classification performance of Gene Expression Data, which projects high dimensional Data onto lower dimensional feature space. We examine a set of Gene Expression Data that include sets of tumor and normal clinical samples by means of SVMs classifier. Experiments show that SVM has a superior performance in classification of Gene Expression Data as long as the selected features can represent the principal components of all Gene Expression samples.

Daxin Jiang - One of the best experts on this subject based on the ideXlab platform.

  • An interactive approach to mining Gene Expression Data
    IEEE Transactions on Knowledge and Data Engineering, 2005
    Co-Authors: Daxin Jiang, Aidong Zhang
    Abstract:

    Effective identification of coexpressed Genes and coherent patterns in Gene Expression Data is an important task in bioinformatics research and biomedical applications. Several clustering methods have recently been proposed to identify coexpressed Genes that share similar coherent patterns. However, there is no objective standard for groups of coexpressed Genes. The interpretation of co-Expression heavily depends on domain knowledge. Furthermore, groups of coexpressed Genes in Gene Expression Data are often highly connected through a large number of "intermediate" Genes. There may be no clear boundaries to separate clusters. Clustering Gene Expression Data also faces the challenges of satisfying biological domain requirements and addressing the high connectivity of the Data sets. In this paper, we propose an interactive framework for exploring coherent patterns in Gene Expression Data. A novel coherent pattern index is proposed to give users highly confident indications of the existence of coherent patterns. To derive a coherent pattern index and facilitate clustering, we devise an attraction tree structure that summarizes the coherence information among Genes in the Data set. We present efficient and scalable algorithms for constructing attraction trees and coherent pattern indices from Gene Expression Data sets. Our experimental results show that our approach is effective in mining Gene Expression Data and is scalable for mining large Data sets.

  • VLDB - GPX: interactive mining of Gene Expression Data
    Proceedings 2004 VLDB Conference, 2004
    Co-Authors: Daxin Jiang, Aidong Zhang
    Abstract:

    Discovering co-expressed Genes and coherent Expression patterns in Gene Expression Data is an important Data analysis task in bioinformatics research and biomedical applications. Although various clustering methods have been proposed, two tough challenges still remain on how to integrate the users' domain knowledge and how to handle the high connectivity in the Data. Recently, we have systematically studied the problem and proposed an effective approach [3]. In this paper, we describe a demonstration of GPX (for Gene Pattern eXplorer), an integrated environment for interactive exploration of coherent Expression patterns and co-expressed Genes in Gene Expression Data. GPX integrates several novel techniques, including the coherent pattern index graph, a Gene annotation panel, and a graphical interface, to adopt users' domain knowledge and support explorative operations in the clustering procedure. The GPX system as well as its techniques will be showcased, and the progress of GPX will be exemplified using several real-world Gene Expression Data sets.

  • Cluster analysis for Gene Expression Data: A survey
    IEEE Transactions on Knowledge and Data Engineering, 2004
    Co-Authors: Daxin Jiang, Chun Tang, Aidong Zhang
    Abstract:

    DNA microarray technology has now made it possible to simultaneously monitor the Expression levels of thousands of Genes during important biological processes and across collections of related samples. Elucidating the patterns hidden in Gene Expression Data offers a tremendous opportunity for an enhanced understanding of functional genomics. However, the large number of Genes and the complexity of biological networks greatly increases the challenges of comprehending and interpreting the resulting mass of Data, which often consists of millions of measurements. A first step toward addressing this challenge is the use of clustering techniques, which is essential in the Data mining process to reveal natural structures and identify interesting patterns in the underlying Data. Cluster analysis seeks to partition a given Data set into groups based on specified features so that the Data points within a group are more similar to each other than the points in different groups. A very rich literature on cluster analysis has developed over the past three decades. Many conventional clustering algorithms have been adapted or directly applied to Gene Expression Data, and also new algorithms have recently been proposed specifically aiming at Gene Expression Data. These clustering algorithms have been proven useful for identifying biologically relevant groups of Genes and samples. In this paper, we first briefly introduce the concepts of microarray technology and discuss the basic elements of clustering on Gene Expression Data. In particular, we divide cluster analysis for Gene Expression Data into three categories. Then, we present specific challenges pertinent to each clustering category and introduce several representative approaches. We also discuss the problem of cluster validation in three aspects and review various methods to assess the quality and reliability of clustering results. Finally, we conclude this paper and suggest the promising trends in this field.