The Experts below are selected from a list of 4341 Experts worldwide ranked by ideXlab platform

Sara C. Madeira - One of the best experts on this subject based on the ideXlab platform.

  • BSig: evaluating the statistical significance of Biclustering solutions
    Data Mining and Knowledge Discovery, 2018
    Co-Authors: Rui Henriques, Sara C. Madeira
    Abstract:

    Statistical evaluation of Biclustering solutions is essential to guarantee the absence of spurious relations and to validate the high number of scientific statements inferred from unsupervised data analysis without a proper statistical ground. Most Biclustering methods rely on merit functions to discover biclusters with specific homogeneity criteria. However, strong homogeneity does not guarantee the statistical significance of Biclustering solutions. Furthermore, although some Biclustering methods test the statistical significance of specific types of biclusters, there are no methods to assess the significance of flexible Biclustering models. This work proposes a method to evaluate the statistical significance of Biclustering solutions. It integrates state-of-the-art statistical views on the significance of local patterns and extends them with new principles to assess the significance of biclusters with additive, multiplicative, symmetric, order-preserving and plaid coherencies. The proposed statistical tests provide the unprecedented possibility to minimize the number of false positive biclusters without incurring on false negatives, and to compare state-of-the-art Biclustering algorithms according to the statistical significance of their outputs. Results on synthetic and real data support the soundness and relevance of the proposed contributions, and stress the need to combine significance and homogeneity criteria to guide the search for biclusters.

  • BicPAMS: software for biological data analysis with pattern-based Biclustering
    BMC Bioinformatics, 2017
    Co-Authors: Rui T Henriques, Francisco L Ferreira, Sara C. Madeira
    Abstract:

    BackgroundBiclustering has been largely applied for the unsupervised analysis of biological data, being recognised today as a key technique to discover putative modules in both expression data (subsets of genes correlated in subsets of conditions) and network data (groups of coherently interconnected biological entities). However, given its computational complexity, only recent breakthroughs on pattern-based Biclustering enabled efficient searches without the restrictions that state-of-the-art Biclustering algorithms place on the structure and homogeneity of biclusters. As a result, pattern-based Biclustering provides the unprecedented opportunity to discover non-trivial yet meaningful biological modules with putative functions, whose coherency and tolerance to noise can be tuned and made problem-specific.MethodsTo enable the effective use of pattern-based Biclustering by the scientific community, we developed BicPAMS (Biclustering based on PAttern Mining Software), a software that: 1) makes available state-of-the-art pattern-based Biclustering algorithms (BicPAM (Henriques and Madeira, Alg Mol Biol 9:27, 2014), BicNET (Henriques and Madeira, Alg Mol Biol 11:23, 2016), BicSPAM (Henriques and Madeira, BMC Bioinforma 15:130, 2014), BiC2PAM (Henriques and Madeira, Alg Mol Biol 11:1–30, 2016), BiP (Henriques and Madeira, IEEE/ACM Trans Comput Biol Bioinforma, 2015), DeBi (Serin and Vingron, AMB 6:1–12, 2011) and BiModule (Okada et al., IPSJ Trans Bioinf 48(SIG5):39–48, 2007)); 2) consistently integrates their dispersed contributions; 3) further explores additional accuracy and efficiency gains; and 4) makes available graphical and application programming interfaces.ResultsResults on both synthetic and real data confirm the relevance of BicPAMS for biological data analysis, highlighting its essential role for the discovery of putative modules with non-trivial yet biologically significant functions from expression and network data.ConclusionsBicPAMS is the first Biclustering tool offering the possibility to: 1) parametrically customize the structure, coherency and quality of biclusters; 2) analyze large-scale biological networks; and 3) tackle the restrictive assumptions placed by state-of-the-art Biclustering algorithms. These contributions are shown to be key for an adequate, complete and user-assisted unsupervised analysis of biological data.SoftwareBicPAMS and its tutorial available in http://www.bicpams.com.

  • BiC2PAM: constraint-guided Biclustering for biological data analysis with domain knowledge
    Algorithms for Molecular Biology, 2016
    Co-Authors: Rui Henriques, Sara C. Madeira
    Abstract:

    BackgroundBiclustering has been largely used in biological data analysis, enabling the discovery of putative functional modules from omic and network data. Despite the recognized importance of incorporating domain knowledge to guide Biclustering and guarantee a focus on relevant and non-trivial biclusters, this possibility has not yet been comprehensively addressed. This results from the fact that the majority of existing algorithms are only able to deliver sub-optimal solutions with restrictive assumptions on the structure, coherency and quality of Biclustering solutions, thus preventing the up-front satisfaction of knowledge-driven constraints. Interestingly, in recent years, a clearer understanding of the synergies between pattern mining and Biclustering gave rise to a new class of algorithms, termed as pattern-based Biclustering algorithms. These algorithms, able to efficiently discover flexible Biclustering solutions with optimality guarantees, are thus positioned as good candidates for knowledge incorporation. In this context, this work aims to bridge the current lack of solid views on the use of background knowledge to guide (pattern-based) Biclustering tasks.MethodsThis work extends (pattern-based) Biclustering algorithms to guarantee the satisfiability of constraints derived from background knowledge and to effectively explore efficiency gains from their incorporation. In this context, we first show the relevance of constraints with succinct, (anti-)monotone and convertible properties for the analysis of expression data and biological networks. We further show how pattern-based Biclustering algorithms can be adapted to effectively prune of the search space in the presence of such constraints, as well as be guided in the presence of biological annotations. Relying on these contributions, we propose Biclustering with Constraints using PAttern Mining (BiC2PAM), an extension of BicPAM and BicNET Biclustering algorithms.ResultsExperimental results on biological data demonstrate the importance of incorporating knowledge within Biclustering to foster efficiency and enable the discovery of non-trivial biclusters with heightened biological relevance.ConclusionsThis work provides the first comprehensive view and sound algorithm for Biclustering biological data with constraints derived from user expectations, knowledge repositories and/or literature.

  • BiC2PAM: constraint-guided Biclustering for biological data analysis with domain knowledge
    Algorithms for molecular biology : AMB, 2016
    Co-Authors: Rui T Henriques, Sara C. Madeira
    Abstract:

    Biclustering has been largely used in biological data analysis, enabling the discovery of putative functional modules from omic and network data. Despite the recognized importance of incorporating domain knowledge to guide Biclustering and guarantee a focus on relevant and non-trivial biclusters, this possibility has not yet been comprehensively addressed. This results from the fact that the majority of existing algorithms are only able to deliver sub-optimal solutions with restrictive assumptions on the structure, coherency and quality of Biclustering solutions, thus preventing the up-front satisfaction of knowledge-driven constraints. Interestingly, in recent years, a clearer understanding of the synergies between pattern mining and Biclustering gave rise to a new class of algorithms, termed as pattern-based Biclustering algorithms. These algorithms, able to efficiently discover flexible Biclustering solutions with optimality guarantees, are thus positioned as good candidates for knowledge incorporation. In this context, this work aims to bridge the current lack of solid views on the use of background knowledge to guide (pattern-based) Biclustering tasks. This work extends (pattern-based) Biclustering algorithms to guarantee the satisfiability of constraints derived from background knowledge and to effectively explore efficiency gains from their incorporation. In this context, we first show the relevance of constraints with succinct, (anti-)monotone and convertible properties for the analysis of expression data and biological networks. We further show how pattern-based Biclustering algorithms can be adapted to effectively prune of the search space in the presence of such constraints, as well as be guided in the presence of biological annotations. Relying on these contributions, we propose Biclustering with Constraints using PAttern Mining (BiC2PAM), an extension of BicPAM and BicNET Biclustering algorithms. Experimental results on biological data demonstrate the importance of incorporating knowledge within Biclustering to foster efficiency and enable the discovery of non-trivial biclusters with heightened biological relevance. This work provides the first comprehensive view and sound algorithm for Biclustering biological data with constraints derived from user expectations, knowledge repositories and/or literature.

  • A structured view on pattern mining-based Biclustering
    Pattern Recognition, 2015
    Co-Authors: Rui T Henriques, Cláudia Antunes, Sara C. Madeira
    Abstract:

    Mining matrices to find relevant biclusters, subsets of rows exhibiting a coherent pattern over a subset of columns, is a critical task for a wide-set of biomedical and social applications. Since Biclustering is a challenging combinatorial optimization task, existing approaches place restrictions on the allowed structure, coherence and quality of biclusters. Biclustering approaches relying on pattern mining (PM) allow an exhaustive yet efficient space exploration together with the possibility to discover flexible structures of biclusters with parameterizable coherency and noise-tolerance. Still, state-of-the-art contributions are dispersed and the potential of their integration remains unclear.This work proposes a structured and integrated view of the contributions of state-of-the-art PM-based Biclustering approaches, makes available a set of principles for a guided definition of new PM-based Biclustering approaches, and discusses their relevance for applications in pattern recognition. Empirical evidence shows that these principles guarantee the robustness, efficiency and flexibility of PM-based Biclustering. HighlightsPattern mining (PM) searches enable flexible, exhaustive and efficient BiclusteringIntegration of existing dispersed PM-inspired contributions for Biclustering.Principles for guided design and evaluation of new PM-based Biclustering approachesPM-based Biclustering solutions have parameterizable coherency and quality.

Patryk Orzechowski - One of the best experts on this subject based on the ideXlab platform.

  • ebic a scalable Biclustering method for large scale data analysis
    Genetic and Evolutionary Computation Conference, 2019
    Co-Authors: Patryk Orzechowski, Jason H. Moore
    Abstract:

    Biclustering is a technique that looks for patterns hidden in some columns and some rows of the input data. Evolutionary search-based Biclustering (EBIC) is probably the first Biclustering method that combines high accuracy of detection of multiple patterns with support for big data. EBIC has been recently extended to a multi-GPU method and allows to analyze very large datasets. In this short paper, we discuss the scalability of EBIC as well as its suitability for RNA-seq and single cell RNA-seq (scRNA-seq) experiments.

  • Scalable Biclustering - the future of big data exploration?
    GigaScience, 2019
    Co-Authors: Patryk Orzechowski, Krzysztof Boryczko, Jason H. Moore
    Abstract:

    Biclustering is a technique of discovering local similarities within data. For many years the complexity of the methods and parallelization issues limited its application to big data problems. With the development of novel scalable methods, Biclustering has finally started to close this gap. In this paper we discuss the caveats of Biclustering and present its current challenges and guidelines for practitioners. We also try to explain why Biclustering may soon become one of the standards for big data analytics.

  • Propagation-Based Biclustering Algorithm for Extracting Inclusion-Maximal Motifs
    Computing and Informatics \ Computers and Artificial Intelligence, 2016
    Co-Authors: Patryk Orzechowski, Krzysztof Boryczko
    Abstract:

    Biclustering, which is simultaneous clustering of columns and rows in data matrix, became an issue when classical clustering algorithms proved not to be good enough to detect similar expressions of genes under subset of conditions. Biclustering algorithms may be also applied to different datasets, such as medical, economical, social networks etc. In this article we explain the concept beneath hybrid Biclustering algorithms and present details of propagation-based Biclustering, a novel approach for extracting inclusion-maximal gene expression motifs conserved in gene microarray data. We prove that this approach may successfully compete with other well-recognized Biclustering algorithms.

  • EvoApplications (1) - Hybrid Biclustering Algorithms for Data Mining
    Applications of Evolutionary Computation, 2016
    Co-Authors: Patryk Orzechowski, Krzysztof Boryczko
    Abstract:

    Hybrid methods are a branch of Biclustering algorithms that emerge from combining selected aspects of pre-existing approaches. The syncretic nature of their construction enriches the existing methods providing them with new properties. In this paper the concept of hybrid Biclustering algorithms is explained. A representative hybrid Biclustering algorithm, inspired by neural networks and associative artificial intelligence, is introduced and the results of its application to microarray data are presented. Finally, the scope and application potential for hybrid Biclustering algorithms is discussed.

  • Hybrid Biclustering Algorithms for Data Mining
    Applications of Evolutionary Computation, 2016
    Co-Authors: Patryk Orzechowski, Krzysztof Boryczko
    Abstract:

    Hybrid methods are a branch of Biclustering algorithms that emerge from combining selected aspects of pre-existing approaches. The syncretic nature of their construction enriches the existing methods providing them with new properties. In this paper the concept of hybrid Biclustering algorithms is explained. A representative hybrid Biclustering algorithm, inspired by neural networks and associative artificial intelligence, is introduced and the results of its application to microarray data are presented. Finally, the scope and application potential for hybrid Biclustering algorithms is discussed.

Hong Yan - One of the best experts on this subject based on the ideXlab platform.

  • A graph spectrum based geometric Biclustering algorithm.
    Journal of theoretical biology, 2012
    Co-Authors: Doris Z. Wang, Hong Yan
    Abstract:

    Biclustering is capable of performing simultaneous clustering on two dimensions of a data matrix and has many applications in pattern classification. For example, in microarray experiments, a subset of genes is co-expressed in a subset of conditions, and Biclustering algorithms can be used to detect the coherent patterns in the data for further analysis of function. In this paper, we present a graph spectrum based geometric Biclustering (GSGBC) algorithm. In the geometrical view, biclusters can be seen as different linear geometrical patterns in high dimensional spaces. Based on this, the modified Hough transform is used to find the Hough vector (HV) corresponding to sub-bicluster patterns in 2D spaces. A graph can be built regarding each HV as a node. The graph spectrum is utilized to identify the eigengroups in which the sub-biclusters are grouped naturally to produce larger biclusters. Through a comparative study, we find that the GSGBC achieves as good a result as GBC and outperforms other kinds of Biclustering algorithms. Also, compared with the original geometrical Biclustering algorithm, it reduces the computing time complexity significantly. We also show that biologically meaningful biclusters can be identified by our method from real microarray gene expression data.

  • Biclustering Analysis for Pattern Discovery: Current Techniques, Comparative Studies and Applications
    Current Bioinformatics, 2012
    Co-Authors: Hongya Zhao, Alan Wee-chung Liew, Doris Z. Wang, Hong Yan
    Abstract:

    Biclustering analysis is a useful methodology to discover the local coherent patterns hidden in a data matrix. Unlike the traditional clustering procedure, which searches for groups of coherent patterns using the entire feature set, Biclustering performs simultaneous pattern classification in both row and column directions in a data matrix. The technique has found useful applications in many fields but notably in bioinformatics. In this paper, we give an overview of the Biclustering problem and review some existing Biclustering algorithms in terms of their underlying methodology, search strategy, detected bicluster patterns, and validation strategies. Moreover, we show that geometry of Biclustering patterns can be used to solve Biclustering problems effectively. Well-known methods in signal and image analysis, such as the Hough transform and relaxation labeling, can be employed to detect the geometrical Biclustering patterns. We present performance evaluation results for several of the well known Biclustering algorithms, on both artificial and real gene expression datasets. Finally, several interesting applications of Biclustering are discussed.

  • a new geometric Biclustering algorithm based on the hough transform for analysis of large scale microarray data
    Journal of Theoretical Biology, 2008
    Co-Authors: Hongya Zhao, Hong Yan, Alan Wee-chung Liew, Xudong Xie
    Abstract:

    Biclustering is an important tool in microarray analysis when only a subset of genes co-regulates in a subset of conditions. Different from standard clustering analyses, Biclustering performs simultaneous classification in both gene and condition directions in a microarray data matrix. However, the Biclustering problem is inherently intractable and computationally complex. In this paper, we present a new Biclustering algorithm based on the geometrical viewpoint of coherent gene expression profiles. In this method, we perform pattern identification based on the Hough transform in a column-pair space. The algorithm is especially suitable for the Biclustering analysis of large-scale microarray data. Our studies show that the approach can discover significant biclusters with respect to the increased noise level and regulatory complexity. Furthermore, we also test the ability of our method to locate biologically verifiable biclusters within an annotated set of genes.

  • FUZZ-IEEE - Fuzzy Biclustering for DNA microarray data analysis
    2008 IEEE International Conference on Fuzzy Systems (IEEE World Congress on Computational Intelligence), 2008
    Co-Authors: Lixin Han, Hong Yan
    Abstract:

    Fuzzy Biclustering analysis is a useful tool for identifying relevant subsets of microarray data. This paper proposes a fuzzy Biclustering clustering method for microarray data analysis. The method employs a combination of the Nelder-Mead and min-max algorithm to construct hierarchically structured Biclustering. The method can automatically identify the groups of genes that show similar expression patterns under a specific subset of the samples.

Jin-kao Hao - One of the best experts on this subject based on the ideXlab platform.

  • Pattern-driven neighborhood search for Biclustering of microarray data
    BMC Bioinformatics, 2012
    Co-Authors: Wassim Ayadi, Mourad Elloumi, Jin-kao Hao
    Abstract:

    Background Biclustering aims at finding subgroups of genes that show highly correlated behaviors across a subgroup of conditions. Biclustering is a very useful tool for mining microarray data and has various practical applications. From a computational point of view, Biclustering is a highly combinatorial search problem and can be solved with optimization methods.

  • iterated local search for Biclustering of microarray data
    Pattern Recognition in Bioinformatics, 2010
    Co-Authors: Wassim Ayadi, Mourad Elloumi, Jin-kao Hao
    Abstract:

    In the context of microarray data analysis, Biclustering aims to identify simultaneously a group of genes that are highly correlated across a group of experimental conditions. This paper presents a Biclustering Iterative Local Search (BILS) algorithm to the problem of Biclustering of microarray data. The proposed algorithm is highlighted by the use of some original features including a new evaluation function, a dedicated neighborhood relation and a tailored perturbation strategy. The BILS algorithm is assessed on the well-known yeast cell-cycle dataset and compared with two most popular algorithms.

  • PRIB - Iterated local search for Biclustering of microarray data
    Pattern Recognition in Bioinformatics, 2010
    Co-Authors: Wassim Ayadi, Mourad Elloumi, Jin-kao Hao
    Abstract:

    In the context of microarray data analysis, Biclustering aims to identify simultaneously a group of genes that are highly correlated across a group of experimental conditions. This paper presents a Biclustering Iterative Local Search (BILS) algorithm to the problem of Biclustering of microarray data. The proposed algorithm is highlighted by the use of some original features including a new evaluation function, a dedicated neighborhood relation and a tailored perturbation strategy. The BILS algorithm is assessed on the well-known yeast cell-cycle dataset and compared with two most popular algorithms.

  • A Biclustering algorithm based on a Bicluster Enumeration Tree: Application to DNA microarray data
    BioData Mining, 2009
    Co-Authors: Wassim Ayadi, Mourad Elloumi, Jin-kao Hao
    Abstract:

    In a number of domains, like in DNA microarray data analysis, we need to cluster simultaneously rows (genes) and columns (conditions) of a data matrix to identify groups of rows coherent with groups of columns. This kind of clustering is called Biclustering. Biclustering algorithms are extensively used in DNA microarray data analysis. More effective Biclustering algorithms are highly desirable and needed.

Jason H. Moore - One of the best experts on this subject based on the ideXlab platform.

  • ebic a scalable Biclustering method for large scale data analysis
    Genetic and Evolutionary Computation Conference, 2019
    Co-Authors: Patryk Orzechowski, Jason H. Moore
    Abstract:

    Biclustering is a technique that looks for patterns hidden in some columns and some rows of the input data. Evolutionary search-based Biclustering (EBIC) is probably the first Biclustering method that combines high accuracy of detection of multiple patterns with support for big data. EBIC has been recently extended to a multi-GPU method and allows to analyze very large datasets. In this short paper, we discuss the scalability of EBIC as well as its suitability for RNA-seq and single cell RNA-seq (scRNA-seq) experiments.

  • Scalable Biclustering - the future of big data exploration?
    GigaScience, 2019
    Co-Authors: Patryk Orzechowski, Krzysztof Boryczko, Jason H. Moore
    Abstract:

    Biclustering is a technique of discovering local similarities within data. For many years the complexity of the methods and parallelization issues limited its application to big data problems. With the development of novel scalable methods, Biclustering has finally started to close this gap. In this paper we discuss the caveats of Biclustering and present its current challenges and guidelines for practitioners. We also try to explain why Biclustering may soon become one of the standards for big data analytics.