The Experts below are selected from a list of 66 Experts worldwide ranked by ideXlab platform

Peter N Robinson - One of the best experts on this subject based on the ideXlab platform.

  • exact score distribution computation for ontological similarity searches
    BMC Bioinformatics, 2011
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/

  • exact score distribution computation for similarity searches in ontologies
    Workshop on Algorithms in Bioinformatics, 2009
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Martin Vingron, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., protein function prediction with the Gene Ontology. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the Human Phenotype Ontology.

Marcel H Schulz - One of the best experts on this subject based on the ideXlab platform.

  • exact score distribution computation for ontological similarity searches
    BMC Bioinformatics, 2011
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/

  • exact score distribution computation for similarity searches in ontologies
    Workshop on Algorithms in Bioinformatics, 2009
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Martin Vingron, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., protein function prediction with the Gene Ontology. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the Human Phenotype Ontology.

Sebastian Bauer - One of the best experts on this subject based on the ideXlab platform.

  • exact score distribution computation for ontological similarity searches
    BMC Bioinformatics, 2011
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/

  • exact score distribution computation for similarity searches in ontologies
    Workshop on Algorithms in Bioinformatics, 2009
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Martin Vingron, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., protein function prediction with the Gene Ontology. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the Human Phenotype Ontology.

Sebastian Kohler - One of the best experts on this subject based on the ideXlab platform.

  • exact score distribution computation for ontological similarity searches
    BMC Bioinformatics, 2011
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., finding functionally related proteins with the Gene Ontology or phenotypically similar diseases with the Human Phenotype Ontology (HPO). We have recently shown that the performance of semantic similarity searches can be improved by ranking results according to the probability of obtaining a given score at random rather than by the scores themselves. However, to date, there are no algorithms for computing the exact distribution of semantic similarity scores, which is necessary for computing the exact P-value of a given score. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the HPO. It is shown that exact P-value calculation improves clinical diagnosis using the HPO compared to approaches based on sampling. The new algorithm enables for the first time exact P-value calculation via exact score distribution computation for ontology similarity searches. The approach is applicable to any ontology for which the annotation-propagation rule holds and can improve any bioinformatic method that makes only use of the raw similarity scores. The algorithm was implemented in Java, supports any ontology in OBO format, and is available for non-commercial and academic usage under: https://compbio.charite.de/svn/hpo/trunk/src/tools/significance/

  • exact score distribution computation for similarity searches in ontologies
    Workshop on Algorithms in Bioinformatics, 2009
    Co-Authors: Marcel H Schulz, Sebastian Kohler, Sebastian Bauer, Martin Vingron, Peter N Robinson
    Abstract:

    Semantic similarity searches in ontologies are an important component of many bioinformatic algorithms, e.g., protein function prediction with the Gene Ontology. In this paper we consider the exact computation of score distributions for similarity searches in ontologies, and introduce a Simple Null Hypothesis which can be used to compute a P-value for the statistical significance of similarity scores. We concentrate on measures based on Resnik's definition of ontological similarity. A new algorithm is proposed that collapses subgraphs of the ontology graph and thereby allows fast score distribution computation. The new algorithm is several orders of magnitude faster than the naive approach, as we demonstrate by computing score distributions for similarity searches in the Human Phenotype Ontology.

Anit Kumar Sahu - One of the best experts on this subject based on the ideXlab platform.

  • Recursive Distributed Detection for Composite Hypothesis Testing: Nonlinear Observation Models in Additive Gaussian Noise
    IEEE Transactions on Information Theory, 2017
    Co-Authors: Anit Kumar Sahu
    Abstract:

    This paper studies recursive composite Hypothesis testing in a network of sparsely connected agents. The network objective is to test a Simple Null Hypothesis against a composite alternative concerning the state of the field, modeled as a vector of (continuous) unknown parameters determining the parametric family of probability measures induced on the agents' observation spaces under the hypotheses. Specifically, under the alternative Hypothesis, each agent sequentially observes an independent and identically distributed time-series consisting of a (nonlinear) function of the true but unknown parameter corrupted by Gaussian noise, whereas, under the Null, they obtain noise only. Two distributed recursive generalized likelihood ratio test type algorithms of the consensus+innovations form are proposed, namely, CIGLRT - L and CIGLRT - NL, in which the agents estimate the underlying parameter and in parallel also update their test decision statistics by simultaneously processing the latest local sensed information and information obtained from neighboring agents. For CIGLRT - NL, for a broad class of nonlinear observation models and under a global observability condition, algorithm parameters which ensure asymptotically decaying probabilities of errors (probability of miss and probability of false detection) are characterized. For CIGLRT - L, a linear observation model is considered and upper bounds on large deviations decay exponent for the error probabilities are obtained.

  • recursive distributed detection for composite Hypothesis testing algorithms and asymptotics
    arXiv: Information Theory, 2016
    Co-Authors: Anit Kumar Sahu
    Abstract:

    This paper studies recursive composite Hypothesis testing in a network of sparsely connected agents. The network objective is to test a Simple Null Hypothesis against a composite alternative concerning the state of the field, modeled as a vector of (continuous) unknown parameters determining the parametric family of probability measures induced on the agents' observation spaces under the hypotheses. Specifically, under the alternative Hypothesis, each agent sequentially observes an independent and identically distributed time-series consisting of a (nonlinear) function of the true but unknown parameter corrupted by Gaussian noise, whereas, under the Null, they obtain noise only. Two distributed recursive generalized likelihood ratio test type algorithms of the \emph{consensus+innovations} form are proposed, namely $\mathcal{CILRT}$ and $\mathcal{CIGLRT}$, in which the agents estimate the underlying parameter and in parallel also update their test decision statistics by simultaneously processing the latest local sensed information and information obtained from neighboring agents. For $\mathcal{CIGLRT}$, for a broad class of nonlinear observation models and under a global observability condition, algorithm parameters which ensure asymptotically decaying probabilities of errors~(probability of miss and probability of false detection) are characterized. For $\mathcal{CILRT}$, a linear observation model is considered and large deviations decay exponents for the error probabilities are obtained.

  • recursive distributed detection for composite Hypothesis testing nonlinear observation models in additive gaussian noise
    arXiv: Information Theory, 2016
    Co-Authors: Anit Kumar Sahu
    Abstract:

    This paper studies recursive composite Hypothesis testing in a network of sparsely connected agents. The network objective is to test a Simple Null Hypothesis against a composite alternative concerning the state of the field, modeled as a vector of (continuous) unknown parameters determining the parametric family of probability measures induced on the agents' observation spaces under the hypotheses. Specifically, under the alternative Hypothesis, each agent sequentially observes an independent and identically distributed time-series consisting of a (nonlinear) function of the true but unknown parameter corrupted by Gaussian noise, whereas, under the Null, they obtain noise only. Two distributed recursive generalized likelihood ratio test type algorithms of the \emph{consensus+innovations} form are proposed, namely $\mathcal{CIGLRT-L}$ and $\mathcal{CIGLRT-NL}$, in which the agents estimate the underlying parameter and in parallel also update their test decision statistics by simultaneously processing the latest local sensed information and information obtained from neighboring agents. For $\mathcal{CIGLRT-NL}$, for a broad class of nonlinear observation models and under a global observability condition, algorithm parameters which ensure asymptotically decaying probabilities of errors~(probability of miss and probability of false detection) are characterized. For $\mathcal{CIGLRT-L}$, a linear observation model is considered and upper bounds on large deviations decay exponent for the error probabilities are obtained.