The Experts below are selected from a list of 67959 Experts worldwide ranked by ideXlab platform
Florence D'alché-buc - One of the best experts on this subject based on the ideXlab platform.
-
Learning a Markov Logic Network for supervised gene regulatory Network Inference
BMC Bioinformatics, 2013Co-Authors: Céline Brouard, Christel Vrain, Julie Dubois, David Castel, Marie-anne Debily, Florence D'alché-bucAbstract:Background Gene regulatory Network Inference remains a challenging problem in systems biology despite the numerous approaches that have been proposed. When substantial knowledge on a gene regulatory Network is already available, supervised Network Inference is appropriate. Such a method builds a binary classifier able to assign a class (Regulation/No regulation) to an ordered pair of genes. Once learnt, the pairwise classifier can be used to predict new regulations. In this work, we explore the framework of Markov Logic Networks (MLN) that combine features of probabilistic graphical models with the expressivity of first-order logic rules. Results We propose to learn a Markov Logic Network, e.g. a set of weighted rules that conclude on the predicate "regulates", starting from a known gene regulatory Network involved in the switch proliferation/differentiation of keratinocyte cells, a set of experimental transcriptomic data and various descriptions of genes all encoded into first-order logic. As training data are unbalanced, we use asymmetric bagging to learn a set of MLNs. The prediction of a new regulation can then be obtained by averaging predictions of individual MLNs. As a side contribution, we propose three in silico tests to assess the performance of any pairwise classifier in various Network Inference tasks on real datasets. A first test consists of measuring the average performance on balanced edge prediction problem; a second one deals with the ability of the classifier, once enhanced by asymmetric bagging, to update a given Network. Finally our main result concerns a third test that measures the ability of the method to predict regulations with a new set of genes. As expected, MLN, when provided with only numerical discretized gene expression data, does not perform as well as a pairwise SVM in terms of AUPR. However, when a more complete description of gene properties is provided by heterogeneous sources, MLN achieves the same performance as a black-box model such as a pairwise SVM while providing relevant insights on the predictions. Conclusions The numerical studies show that MLN achieves very good predictive performance while opening the door to some interpretability of the decisions. Besides the ability to suggest new regulations, such an approach allows to cross-validate experimental data with existing knowledge.
-
Boosting an operator-valued kernel-based model for gene regulatory Network Inference
2013Co-Authors: Néhémy Lim, George Michailidis, Yasin Senbabaoglu, Florence D'alché-bucAbstract:Motivation: Reverse engineering of gene regulatory Networks remains a central challenge in computational systems biology, despite recent advances facilitated by benchmark in-silico challenges that have aided in calibrating their performance. A number of approaches using either perturbation (knock-out) or wild-type time series data have appeared in the literature addressing this problem, with the latter employing linear temporal models. Nonlinear dynamical models are particularly appropriate for this Inference task given the generation mechanism of the time series data. In this study, we introduce a novel nonlinear autoregressive model based on operator-valued kernels that simultaneously learns the model parameters, as well as the Network structure. Model and Methods: A flexible boosting algorithm (OKVAR-Boost) that shares features from L2-boosting and randomization-based algorithms is developed to perform the tasks of parameter learning and Network Inference for the proposed model. Specifically, at each boosting iteration, a regularized operator-valued kernel based vector autoregressive model (OKVAR) is trained on a random subNetwork. The final model consists of an ensemble of such models. The empirical estimation of the ensemble model's Jacobian matrix provides an estimation of the Network structure. Results:This study makes a number of key contributions to the challenging problem of Network Inference based solely on time course data. It introduces a powerful Network Inference framework based on nonlinear autoregressive modeling and Jacobian estimation. The proposed framework is rich and flexible, employing penalized regression models that coupled with randomized search algorithms and features of L2-boosting prove particularly effective as the extensive simulation results attest. The models employed require tuning of a number of parameters and we introduce a novel and generally applicable strategy that combines bootstrapping with stability selection to achieve this goal. The performance of the proposed algorithm is first evaluated on a number of benchmark data sets from the DREAM3 challenge and then, on real datasets related to the IRMA and T-cell Networks. The high quality results obtained strongly indicate that it outperforms existing approaches.
-
Boosting an operator-valued kernel model and application to Network Inference
2013Co-Authors: Néhémy Lim, George Michailidis, Yasin Senbabaoglu, Florence D'alché-bucAbstract:Reverse engineering of gene regulatory Networks remains a central challenge in computational systems biology, despite recent advances facilitated by benchmark in-silico challenges that have aided in calibrating their performance. A number of approaches using either perturbation (knock-out) or wild-type time series data have appeared in the literature addressing this problem, with the latter employing linear temporal models. Nonlinear dynamical models are particularly appropriate for this Inference task given the generation mechanism of the time series data. In this study, we introduce a novel nonlinear autoregressive model based on operator-valued kernels that simultaneously learns the model parameters, as well as the Network structure. A flexible boosting algorithm (OKVAR-Boost) that shares features from L2-boosting and randomization-based algorithms is developed to perform the tasks of parameter learning and Network Inference for the proposed model. Speci cally, at each boosting iteration, a regularized operator-valued kernel based vector autoregressive model (OKVAR) is trained on a random sub-Network. The fi nal model consists of an ensemble of such models. The empirical estimation of the ensemble model's Jacobian matrix provides an estimation of the Network structure. The performance of the proposed algorithm is evaluated on a number of benchmark data sets from the DREAM3 challenge. The high quality results obtained strongly indicate that it outperforms existing approaches.
-
Protein-protein interaction Network Inference with semi-supervised Output Kernel Regression
2012Co-Authors: Céline Brouard, Marie Szafranski, Florence D'alché-bucAbstract:In this work, we address the problem of protein-protein interaction Network Inference as a semi-supervised output kernel learning problem. Using the kernel trick in the output space allows one to reduce the problem of learning from pairs to learning a single variable function with values in a Hilbert space. We turn to the Reproducing Kernel Hilbert Space theory devoted to vector- valued functions, which provides us with a general framework for output kernel regression. In this framework, we propose a novel method which allows to extend Output Kernel Regression to semi-supervised learning. We study the relevance of this approach on transductive link prediction using artificial data and a protein-protein interaction Network of S. Cerevisiae using a very low percentage of labeled data.
-
Estimation of nonparametric dynamical models within Reproducing Kernel Hilbert Spaces for Network Inference
2012Co-Authors: Florence D'alché-buc, Néhémy Lim, George Michailidis, Yasin SenbabaogluAbstract:We consider the problem of Network Inference that occurs for instance in systems biology. A dynamical system (a gene regulatory Network) is observed through time and the goal is to infer the dependence structure between state variables (mRNAs concentrations) from time series. Works concerning net- work Inference usually rely on sparse linear models estimation or Granger causality tools. A very few address the issue in the nonlinear cases. In this work, we propose a nonparametric approach to dynamical system modeling that makes no assumption about the nature of the underlying nonlinear system. We develop a general framework based on Reproducing Kernel Hilbert Spaces based on matrix-valued kernels to identify the dynamical system and retrieve the target Network. As in the linear case, the Network Inference task calls for sparsity control. We show very good results both in autoregressive models and differential equations estimation on DREAM benchmarks as well as on the IRMA datasets.
Céline Brouard - One of the best experts on this subject based on the ideXlab platform.
-
Learning a Markov Logic Network for supervised gene regulatory Network Inference
BMC Bioinformatics, 2013Co-Authors: Céline Brouard, Christel Vrain, Julie Dubois, David Castel, Marie-anne Debily, Florence D'alché-bucAbstract:Background Gene regulatory Network Inference remains a challenging problem in systems biology despite the numerous approaches that have been proposed. When substantial knowledge on a gene regulatory Network is already available, supervised Network Inference is appropriate. Such a method builds a binary classifier able to assign a class (Regulation/No regulation) to an ordered pair of genes. Once learnt, the pairwise classifier can be used to predict new regulations. In this work, we explore the framework of Markov Logic Networks (MLN) that combine features of probabilistic graphical models with the expressivity of first-order logic rules. Results We propose to learn a Markov Logic Network, e.g. a set of weighted rules that conclude on the predicate "regulates", starting from a known gene regulatory Network involved in the switch proliferation/differentiation of keratinocyte cells, a set of experimental transcriptomic data and various descriptions of genes all encoded into first-order logic. As training data are unbalanced, we use asymmetric bagging to learn a set of MLNs. The prediction of a new regulation can then be obtained by averaging predictions of individual MLNs. As a side contribution, we propose three in silico tests to assess the performance of any pairwise classifier in various Network Inference tasks on real datasets. A first test consists of measuring the average performance on balanced edge prediction problem; a second one deals with the ability of the classifier, once enhanced by asymmetric bagging, to update a given Network. Finally our main result concerns a third test that measures the ability of the method to predict regulations with a new set of genes. As expected, MLN, when provided with only numerical discretized gene expression data, does not perform as well as a pairwise SVM in terms of AUPR. However, when a more complete description of gene properties is provided by heterogeneous sources, MLN achieves the same performance as a black-box model such as a pairwise SVM while providing relevant insights on the predictions. Conclusions The numerical studies show that MLN achieves very good predictive performance while opening the door to some interpretability of the decisions. Besides the ability to suggest new regulations, such an approach allows to cross-validate experimental data with existing knowledge.
-
Protein-protein interaction Network Inference with semi-supervised Output Kernel Regression
2012Co-Authors: Céline Brouard, Marie Szafranski, Florence D'alché-bucAbstract:In this work, we address the problem of protein-protein interaction Network Inference as a semi-supervised output kernel learning problem. Using the kernel trick in the output space allows one to reduce the problem of learning from pairs to learning a single variable function with values in a Hilbert space. We turn to the Reproducing Kernel Hilbert Space theory devoted to vector- valued functions, which provides us with a general framework for output kernel regression. In this framework, we propose a novel method which allows to extend Output Kernel Regression to semi-supervised learning. We study the relevance of this approach on transductive link prediction using artificial data and a protein-protein interaction Network of S. Cerevisiae using a very low percentage of labeled data.
-
A new theoretical angle to semi-supervised output kernel regression for protein-protein interaction Network Inference
2011Co-Authors: Céline Brouard, Florence D'alché-buc, Marie SzafranskiAbstract:Protein-protein interaction Network Inference is addressed as an output kernel learning task through semi-supervised Output Kernel Regression. Working in the framework of RKHS theory for vector-valued functions, we establish a new representer theorem devoted to semi-supervised least square regression. We then apply it to get a new model and show its relevance using numerical experiments on artificial Networks and a protein-protein interaction Network dataset using a very low percentage of labeled proteins in a transductive setting.
Michael P. H. Stumpf - One of the best experts on this subject based on the ideXlab platform.
-
Gene regulatory Network Inference
Systems Medicine, 2021Co-Authors: Ann C. Babtie, Michael P. H. Stumpf, Thomas ThorneAbstract:Transcriptomic data quantifying gene expression states for single cells or cell populations at a genomic level is now readily available. Changes in transcriptional state during cell development and function are governed by gene regulatory Networks, comprising a collection of genes and regulatory interactions between these genes (or gene products). Network Inference algorithms aim to infer functional interactions between genes from experimentally observed expression profiles, and identify the structure of the underlying regulatory Networks. Here we describe popular classes of Network Inference algorithms, highlighting their respective strengths and weaknesses, along with some general challenges faced by these methods. Analyzing inferred Network structures can provide insight into the genes, transcriptional changes, and regulatory interactions that play key roles in biological and disease-related processes of interest.
-
Parametric and non-parametric gradient matching for Network Inference: a comparison
BMC Bioinformatics, 2019Co-Authors: Leander Dony, Michael P. H. StumpfAbstract:Background Reverse engineering of gene regulatory Networks from time series gene-expression data is a challenging problem, not only because of the vast sets of candidate interactions but also due to the stochastic nature of gene expression. We limit our analysis to nonlinear differential equation based Inference methods. In order to avoid the computational cost of large-scale simulations, a two-step Gaussian process interpolation based gradient matching approach has been proposed to solve differential equations approximately. Results We apply a gradient matching Inference approach to a large number of candidate models, including parametric differential equations or their corresponding non-parametric representations, we evaluate the Network Inference performance under various settings for different Inference objectives. We use model averaging, based on the Bayesian Information Criterion (BIC), to combine the different Inferences. The performance of different Inference approaches is evaluated using area under the precision-recall curves. Conclusions We found that parametric methods can provide comparable, and often improved Inference compared to non-parametric methods; the latter, however, require no kinetic information and are computationally more efficient.
-
gene regulatory Network Inference from single cell data using multivariate information measures
Cell systems, 2017Co-Authors: Thalia E Chan, Michael P. H. Stumpf, Ann C. BabtieAbstract:While single-cell gene expression experiments present new challenges for data processing, the cell-to-cell variability observed also reveals statistical relationships that can be used by information theory. Here, we use multivariate information theory to explore the statistical dependencies between triplets of genes in single-cell gene expression datasets. We develop PIDC, a fast, efficient algorithm that uses partial information decomposition (PID) to identify regulatory relationships between genes. We thoroughly evaluate the performance of our algorithm and demonstrate that the higher-order information captured by PIDC allows it to outperform pairwise mutual information-based algorithms when recovering true relationships present in simulated data. We also infer gene regulatory Networks from three experimental single-cell datasets and illustrate how Network context, choices made during analysis, and sources of variability affect Network Inference. PIDC tutorials and open-source software for estimating PID are available. PIDC should facilitate the identification of putative functional relationships and mechanistic hypotheses from single-cell transcriptomic data.
-
gene regulatory Network Inference from single cell data using multivariate information measures
bioRxiv, 2017Co-Authors: Thalia E Chan, Michael P. H. Stumpf, Ann C. BabtieAbstract:While single-cell gene expression experiments present new challenges for data processing, the cell-to-cell variability observed also reveals statistical relationships that can be used by information theory. Here, we use multivariate information theory to explore the statistical dependencies between triplets of genes in single-cell gene expression datasets. We develop PIDC, a fast, efficient algorithm that uses partial information decomposition (PID) to identify regulatory relationships between genes. We thoroughly evaluate the performance of our algorithm and demonstrate that the higher order information captured by PIDC allows it to outperform pairwise mutual information-based algorithms when recovering true relationships present in simulated data. We also infer gene regulatory Networks from three experimental single-cell data sets and illustrate how Network context, choices made during analysis, and sources of variability affect Network Inference. PIDC tutorials and open-source software for estimating PID are available here: https://github.com/Tchanders/Network_Inference_tutorials. PIDC should facilitate the identification of putative functional relationships and mechanistic hypotheses from single-cell transcriptomic data.
Yungkeun Kwon - One of the best experts on this subject based on the ideXlab platform.
-
a boolean Network Inference from time series gene expression data using a genetic algorithm
Bioinformatics, 2018Co-Authors: Shohag Barman, Yungkeun KwonAbstract:Motivation Inferring a gene regulatory Network from time-series gene expression data is a fundamental problem in systems biology, and many methods have been proposed. However, most of them were not efficient in inferring regulatory relations involved by a large number of genes because they limited the number of regulatory genes or computed an approximated reliability of multivariate relations. Therefore, an improved method is needed to efficiently search more generalized and scalable regulatory relations. Results In this study, we propose a genetic algorithm-based Boolean Network Inference (GABNI) method which can search an optimal Boolean regulatory function of a large number of regulatory genes. For an efficient search, it solves the problem in two stages. GABNI first exploits an existing method, a mutual information-based Boolean Network Inference (MIBNI), because it can quickly find an optimal solution in a small-scale Inference problem. When MIBNI fails to find an optimal solution, a genetic algorithm (GA) is applied to search an optimal set of regulatory genes in a wider solution space. In particular, we modified a typical GA framework to efficiently reduce a search space. We compared GABNI with four well-known Inference methods through extensive simulations on both the artificial and the real gene expression datasets. Our results demonstrated that GABNI significantly outperformed them in both structural and dynamics accuracies. Conclusion The proposed method is an efficient and scalable tool to infer a Boolean Network from time-series gene expression data. Supplementary information Supplementary data are available at Bioinformatics online.
-
a novel mutual information based boolean Network Inference method from time series gene expression data
PLOS ONE, 2017Co-Authors: Shohag Barman, Yungkeun KwonAbstract:Background Inferring a gene regulatory Network from time-series gene expression data in systems biology is a challenging problem. Many methods have been suggested, most of which have a scalability limitation due to the combinatorial cost of searching a regulatory set of genes. In addition, they have focused on the accurate Inference of a Network structure only. Therefore, there is a pressing need to develop a Network Inference method to search regulatory genes efficiently and to predict the Network dynamics accurately. Results In this study, we employed a Boolean Network model with a restricted update rule scheme to capture coarse-grained dynamics, and propose a novel mutual information-based Boolean Network Inference (MIBNI) method. Given time-series gene expression data as an input, the method first identifies a set of initial regulatory genes using mutual information-based feature selection, and then improves the dynamics prediction accuracy by iteratively swapping a pair of genes between sets of the selected regulatory genes and the other genes. Through extensive simulations with artificial datasets, MIBNI showed consistently better performance than six well-known existing methods, REVEAL, Best-Fit, RelNet, CST, CLR, and BIBN in terms of both structural and dynamics prediction accuracy. We further tested the proposed method with two real gene expression datasets for an Escherichia coli gene regulatory Network and a fission yeast cell cycle Network, and also observed better results using MIBNI compared to the six other methods. Conclusions Taken together, MIBNI is a promising tool for predicting both the structure and the dynamics of a gene regulatory Network.
Arnaud Bonnaffoux - One of the best experts on this subject based on the ideXlab platform.
-
wasabi a dynamic iterative framework for gene regulatory Network Inference
BMC Bioinformatics, 2019Co-Authors: Arnaud Bonnaffoux, Ulysse Herbach, Angelique Richard, Anissa Guillemin, Sandrine Goningiraud, Pierrealexis Gros, Olivier GandrillonAbstract:Background Inference of gene regulatory Networks from gene expression data has been a long-standing and notoriously difficult task in systems biology. Recently, single-cell transcriptomic data have been massively used for gene regulatory Network Inference, with both successes and limitations.
-
WASABI: a dynamic iterative framework for gene regulatory Network Inference
BMC Bioinformatics, 2019Co-Authors: Arnaud Bonnaffoux, Ulysse Herbach, Angelique Richard, Anissa Guillemin, Pierrealexis Gros, Sandrine Gonin-giraud, Olivier GandrillonAbstract:Background Inference of gene regulatory Networks from gene expression data has been a long-standing and notoriously difficult task in systems biology. Recently, single-cell transcriptomic data have been massively used for gene regulatory Network Inference, with both successes and limitations. Results In the present work we propose an iterative algorithm called WASABI, dedicated to inferring a causal dynamical Network from time-stamped single-cell data, which tackles some of the limitations associated with current approaches. We first introduce the concept of waves, which posits that the information provided by an external stimulus will affect genes one-by-one through a cascade, like waves spreading through a Network. This concept allows us to infer the Network one gene at a time, after genes have been ordered regarding their time of regulation. We then demonstrate the ability of WASABI to correctly infer small Networks, which have been simulated in silico using a mechanistic model consisting of coupled piecewise-deterministic Markov processes for the proper description of gene expression at the single-cell level. We finally apply WASABI on in vitro generated data on an avian model of erythroid differentiation. The structure of the resulting gene regulatory Network sheds a new light on the molecular mechanisms controlling this process. In particular, we find no evidence for hub genes and a much more distributed Network structure than expected. Interestingly, we find that a majority of genes are under the direct control of the differentiation-inducing stimulus. Conclusions Together, these results demonstrate WASABI versatility and ability to tackle some general gene regulatory Networks Inference issues. It is our hope that WASABI will prove useful in helping biologists to fully exploit the power of time-stamped single-cell data.
-
wasabi a dynamic iterative framework for gene regulatory Network Inference
bioRxiv, 2018Co-Authors: Arnaud Bonnaffoux, Ulysse Herbach, Angelique Richard, Anissa Guillemin, Pierrealexis Gros, Sandrine Giraud, Olivier GandrillonAbstract:Abstract Inference of gene regulatory Networks from gene expression data has been a long-standing and notoriously difficult task in systems biology. Recently, single-cell transcriptomic data have been massively used for gene regulatory Network Inference, with both successes and limitations. In the present work we propose an iterative algorithm called WASABI, dedicated to inferring a causal dynamical Network from time-stamped single-cell data, which tackles some of the limitations associated with current approaches. We first introduce the concept of waves, which posits that the information provided by an external stimulus will affect genes one-by-one through a cascade, like waves spreading through a Network. This concept allows us to infer the Network one gene at a time, after genes have been ordered regarding their time of regulation. We then demonstrate the ability of WASABI to correctly infer small Networks, which have been simulated in silico using a mechanistic model consisting of coupled piecewise-deterministic Markov processes for the proper description of gene expression at the single-cell level. We finally apply WASABI on in vitro generated data on an avian model of erythroid differentiation. The structure of the resulting gene regulatory Network sheds a fascinating new light on the molecular mechanisms controlling this process. In particular, we find no evidence for hub genes and a much more distributed Network structure than expected. Interestingly, we find that a majority of genes are under the direct control of the differentiation-inducing stimulus. In conclusion, WASABI is a versatile algorithm which should help biologists to fully exploit the power of time-stamped single-cell data.