The Experts below are selected from a list of 975 Experts worldwide ranked by ideXlab platform

Ola Spjuth - One of the best experts on this subject based on the ideXlab platform.

  • predicting off target binding profiles with confidence using Conformal Prediction
    Frontiers in Pharmacology, 2018
    Co-Authors: Samuel Lampa, Ernst Ahlberg, Jonathan Alvarsson, Staffan Arvidsson Mc Shane, Arvid Berg, Ola Spjuth
    Abstract:

    Ligand-based models can be used in drug discovery to obtain an early indication of potential off-target interactions that could be linked to adverse effects. Another application is to combine such models into a panel, allowing to compare and search for compounds with similar profiles. Most contemporary methods and implementations however lack valid measures of confidence in their Predictions, and only provide point Predictions. We here describe a methodology that uses Conformal Prediction for predicting off-target interactions, with models trained on data from 31 targets in the ExCAPE-DB dataset selected for their utility in broad early hazard assessment. Chemicals were represented by the signature molecular descriptor and support vector machines were used as the underlying machine learning method. By using Conformal Prediction, the results from Predictions come in the form of confidence p-values for each class. The full pre-processing and model training process is openly available as scientific workflows on GitHub, rendering it fully reproducible. We illustrate the usefulness of the developed methodology on a set of compounds extracted from DrugBank. The resulting models are published online and are available via a graphical web interface and an OpenAPI interface for programmatic access.

  • Aggregating Predictions on Multiple Non-disclosed Datasets using Conformal Prediction.
    arXiv: Machine Learning, 2018
    Co-Authors: Ola Spjuth, Lars Carlsson, Niharika Gauraha
    Abstract:

    Conformal Prediction is a machine learning methodology that produces valid Prediction regions under mild conditions. In this paper, we explore the application of making Predictions over multiple data sources of different sizes without disclosing data between the sources. We propose that each data source applies a transductive Conformal predictor independently using the local data, and that the individual Predictions are then aggregated to form a combined Prediction region. We demonstrate the method on several data sets, and show that the proposed method produces conservatively valid Predictions and reduces the variance in the aggregated Predictions. We also study the effect that the number of data sources and size of each source has on aggregated Predictions, as compared with equally sized sources and pooled data.

  • ConformalClassification: A Conformal Prediction R Package for Classification
    arXiv: Machine Learning, 2018
    Co-Authors: Niharika Gauraha, Ola Spjuth
    Abstract:

    The ConformalClassification package implements Transductive Conformal Prediction (TCP) and Inductive Conformal Prediction (ICP) for classification problems. Conformal Prediction (CP) is a framework that complements the Predictions of machine learning algorithms with reliable measures of confidence. TCP gives results with higher validity than ICP, however ICP is computationally faster than TCP. The package ConformalClassification is built upon the random forest method, where votes of the random forest for each class are considered as the conformity scores for each data point. Although the main aim of the ConformalClassification package is to generate CP errors (p-values) for classification problems, the package also implements various diagnostic measures such as deviation from validity, error rate, efficiency, observed fuzziness and calibration plots. In future releases, we plan to extend the package to use other machine learning algorithms, (e.g. support vector machines) for model fitting.

  • Conformal Prediction in learning under privileged information paradigm with applications in drug discovery
    arXiv: Machine Learning, 2018
    Co-Authors: Niharika Gauraha, Lars Carlsson, Ola Spjuth
    Abstract:

    This paper explores Conformal Prediction in the learning under privileged information (LUPI) paradigm. We use the SVM+ realization of LUPI in an inductive Conformal predictor, and apply it to the MNIST benchmark dataset and three datasets in drug discovery. The results show that using privileged information produces valid models and improves efficiency compared to standard SVM, however the improvement varies between the tested datasets and is not substantial in the drug discovery applications. More importantly, using SVM+ in a Conformal Prediction framework enables valid Prediction intervals at specified significance levels.

  • Efficient iterative virtual screening with Apache Spark and Conformal Prediction
    Journal of Cheminformatics, 2018
    Co-Authors: Laeeq Ahmed, Marco Capuccini, Valentin Georgiev, Salman Toor, Wesley Schaal, Erwin Laure, Ola Spjuth
    Abstract:

    Background Docking and scoring large libraries of ligands against target proteins forms the basis of structure-based virtual screening. The problem is trivially parallelizable, and calculations are generally carried out on computer clusters or on large workstations in a brute force manner, by docking and scoring all available ligands. Contribution In this study we propose a strategy that is based on iteratively docking a set of ligands to form a training set, training a ligand-based model on this set, and predicting the remainder of the ligands to exclude those predicted as ‘low-scoring’ ligands. Then, another set of ligands are docked, the model is retrained and the process is repeated until a certain model efficiency level is reached. Thereafter, the remaining ligands are docked or excluded based on this model. We use SVM and Conformal Prediction to deliver valid Prediction intervals for ranking the predicted ligands, and Apache Spark to parallelize both the docking and the modeling. Results We show on 4 different targets that Conformal Prediction based virtual screening (CPVS) is able to reduce the number of docked molecules by 62.61% while retaining an accuracy for the top 30 hits of 94% on average and a speedup of 3.7. The implementation is available as open source via GitHub ( https://github.com/laeeq80/spark-cpvs ) and can be run on high-performance computers as well as on cloud resources.

Ulf Norinder - One of the best experts on this subject based on the ideXlab platform.

  • Conformal Prediction of hdac inhibitors
    Sar and Qsar in Environmental Research, 2019
    Co-Authors: Ulf Norinder, Jesus J Naveja, Edgar Lopezlopez, Daniel Mucs, Jose L Medinafranco
    Abstract:

    : The growing interest in epigenetic probes and drug discovery, as revealed by several epigenetic drugs in clinical use or in the lineup of the drug development pipeline, is boosting the generation of screening data. In order to maximize the use of structure-activity relationships there is a clear need to develop robust and accurate models to understand the underlying structure-activity relationship. Similarly, accurate models should be able to guide the rational screening of compound libraries. Herein we introduce a novel approach for epigenetic quantitative structure-activity relationship (QSAR) modelling using Conformal Prediction. As a case study, we discuss the development of models for 11 sets of inhibitors of histone deacetylases (HDACs), which are one of the major epigenetic target families that have been screened. It was found that all derived models, for every HDAC endpoint and all three significance levels, are valid with respect to Predictions for the external test sets as well as the internal validation of the corresponding training sets. Furthermore, the efficiencies for the Predictions are above 80% for most data sets and above 90% for four data sets at different significant levels. The findings of this work encourage prospective applications of Conformal Prediction for other epigenetic target data sets.

  • multitask modeling with confidence using matrix factorization and Conformal Prediction
    Journal of Chemical Information and Modeling, 2019
    Co-Authors: Ulf Norinder, Fredrik Svensson
    Abstract:

    Multitask Prediction of bioactivities is often faced with challenges relating to the sparsity of data and imbalance between different labels. We propose class conditional (Mondrian) Conformal predictors using underlying Macau models as a novel approach for large scale bioactivity Prediction. This approach handles both high degrees of missing data and label imbalances while still producing high quality predictive models. When applied to ten assay end points from PubChem, the models generated valid models with an efficiency of 74.0–80.1% at the 80% confidence level with similar performance both for the minority and majority class. Also when deleting progressively larger portions of the available data (0–80%) the performance of the models remained robust with only minor deterioration (reduction in efficiency between 5 and 10%). Compared to using Macau without Conformal Prediction the method presented here significantly improves the performance on imbalanced data sets.

  • predicting ames mutagenicity using Conformal Prediction in the ames qsar international challenge project
    Mutagenesis, 2019
    Co-Authors: Ernst Ahlberg, Ulf Norinder, Lars Carlsson
    Abstract:

    : Valid and predictive models for classifying Ames mutagenicity have been developed using Conformal Prediction. The models are Random Forest models using signature molecular descriptors. The investigation indicates, on excluding not-strongly mutagenic compounds (class B), that the validity for mutagenic compounds is increased for the Predictions based on both public and the Division of Genetics and Mutagenesis, National Institute of Health Sciences of Japan (DGM/NIHS) data while less so when using only the latter data source. The former models only result in valid Predictions for the majority, non-mutagenic, class whereas the latter models are valid for both classes, i.e. mutagenic and non-mutagenic compounds. These results demonstrate the importance of data consistency manifested through the superior predictive quality and validity of the models based only on DGM/NIHS generated data compared to a combination of this data with public data sources.

  • Predicting Ames Mutagenicity Using Conformal Prediction in the Ames/QSAR International Challenge Project.
    Mutagenesis, 2018
    Co-Authors: Ulf Norinder, Ernst Ahlberg, Lars Carlsson
    Abstract:

    : Valid and predictive models for classifying Ames mutagenicity have been developed using Conformal Prediction. The models are Random Forest models using signature molecular descriptors. The investigation indicates, on excluding not-strongly mutagenic compounds (class B), that the validity for mutagenic compounds is increased for the Predictions based on both public and the Division of Genetics and Mutagenesis, National Institute of Health Sciences of Japan (DGM/NIHS) data while less so when using only the latter data source. The former models only result in valid Predictions for the majority, non-mutagenic, class whereas the latter models are valid for both classes, i.e. mutagenic and non-mutagenic compounds. These results demonstrate the importance of data consistency manifested through the superior predictive quality and validity of the models based only on DGM/NIHS generated data compared to a combination of this data with public data sources.

  • predicting aromatic amine mutagenicity with confidence a case study using Conformal Prediction
    Biomolecules, 2018
    Co-Authors: Ulf Norinder, Glenn Myatt, Ernst Ahlberg
    Abstract:

    The occurrence of mutagenicity in primary aromatic amines has been investigated using Conformal Prediction. The results of the investigation show that it is possible to develop mathematically proven valid models using Conformal Prediction and that the existence of uncertain classes of Prediction, such as both (both classes assigned to a compound) and empty (no class assigned to a compound), provides the user with additional information on how to use, further develop, and possibly improve future models. The study also indicates that the use of different sets of fingerprints results in models, for which the ability to discriminate varies with respect to the set level of acceptable errors.

Harris Papadopoulos - One of the best experts on this subject based on the ideXlab platform.

  • cross Conformal Prediction with ridge regression
    International Symposium on Statistical Learning and Data Sciences, 2015
    Co-Authors: Harris Papadopoulos
    Abstract:

    Cross-Conformal Prediction (CCP) is a recently proposed approach for overcoming the computational inefficiency problem of Conformal Prediction (CP) without sacrificing as much informational efficiency as Inductive Conformal Prediction (ICP). In effect CCP is a hybrid approach combining the ideas of cross-validation and ICP. In the case of classification the Predictions of CCP have been shown to be empirically valid and more informationally efficient than those of the ICP. This paper introduces CCP in the regression setting and examines its empirical validity and informational efficiency compared to that of the original CP and ICP when combined with Ridge Regression.

  • SLDS - Cross-Conformal Prediction with Ridge Regression
    Statistical Learning and Data Sciences, 2015
    Co-Authors: Harris Papadopoulos
    Abstract:

    Cross-Conformal Prediction (CCP) is a recently proposed approach for overcoming the computational inefficiency problem of Conformal Prediction (CP) without sacrificing as much informational efficiency as Inductive Conformal Prediction (ICP). In effect CCP is a hybrid approach combining the ideas of cross-validation and ICP. In the case of classification the Predictions of CCP have been shown to be empirically valid and more informationally efficient than those of the ICP. This paper introduces CCP in the regression setting and examines its empirical validity and informational efficiency compared to that of the original CP and ICP when combined with Ridge Regression.

  • regression Conformal Prediction with nearest neighbours
    Journal of Artificial Intelligence Research, 2011
    Co-Authors: Harris Papadopoulos, Vladimir Vovk, Alexander Gammerman
    Abstract:

    In this paper we apply Conformal Prediction (CP) to the k-Nearest Neighbours Regression (k-NNR) algorithm and propose ways of extending the typical nonconformity measure used for regression so far. Unlike traditional regression methods which produce point Predictions, Conformal Predictors output predictive regions that satisfy a given confidence level. The regions produced by any Conformal Predictor are automatically valid, however their tightness and therefore usefulness depends on the nonconformity measure used by each CP. In effect a nonconformity measure evaluates how strange a given example is compared to a set of other examples based on some traditional machine learning algorithm. We define six novel nonconformity measures based on the k-Nearest Neighbours Regression algorithm and develop the corresponding CPs following both the original (transductive) and the inductive CP approaches. A comparison of the predictive regions produced by our measures with those of the typical regression measure suggests that a major improvement in terms of predictive region tightness is achieved by the new measures.

  • assessment of stroke risk based on morphological ultrasound image analysis with Conformal Prediction
    Artificial Intelligence Applications and Innovations, 2010
    Co-Authors: Antonis Lambrou, Harris Papadopoulos, Alexander Gammerman, Efthyvoulos Kyriacou, C S Pattichis, Marios S Pattichis, A N Nicolaides
    Abstract:

    Non-invasive ultrasound imaging of carotid plaques allows for the development of plaque image analysis in order to assess the risk of stroke. In our work, we provide reliable confidence measures for the assessment of stroke risk, using the Conformal Prediction framework. This framework provides a way for assigning valid confidence measures to Predictions of classical machine learning algorithms. We conduct experiments on a dataset which contains morphological features derived from ultrasound images of atherosclerotic carotid plaques, and we evaluate the results of four different Conformal Predictors (CPs). The four CPs are based on Artificial Neural Networks (ANNs), Support Vector Machines (SVMs), Naive Bayes classification (NBC), and k-Nearest Neighbours (k-NN). The results given by all CPs demonstrate the reliability and usefulness of the obtained confidence measures on the problem of stroke risk assessment.

  • AIAI - Assessment of Stroke Risk Based on Morphological Ultrasound Image Analysis with Conformal Prediction
    IFIP Advances in Information and Communication Technology, 2010
    Co-Authors: Antonis Lambrou, Harris Papadopoulos, Alexander Gammerman, Efthyvoulos Kyriacou, C S Pattichis, Marios S Pattichis, A N Nicolaides
    Abstract:

    Non-invasive ultrasound imaging of carotid plaques allows for the development of plaque image analysis in order to assess the risk of stroke. In our work, we provide reliable confidence measures for the assessment of stroke risk, using the Conformal Prediction framework. This framework provides a way for assigning valid confidence measures to Predictions of classical machine learning algorithms. We conduct experiments on a dataset which contains morphological features derived from ultrasound images of atherosclerotic carotid plaques, and we evaluate the results of four different Conformal Predictors (CPs). The four CPs are based on Artificial Neural Networks (ANNs), Support Vector Machines (SVMs), Naive Bayes classification (NBC), and k-Nearest Neighbours (k-NN). The results given by all CPs demonstrate the reliability and usefulness of the obtained confidence measures on the problem of stroke risk assessment.

Lars Carlsson - One of the best experts on this subject based on the ideXlab platform.

  • predicting ames mutagenicity using Conformal Prediction in the ames qsar international challenge project
    Mutagenesis, 2019
    Co-Authors: Ernst Ahlberg, Ulf Norinder, Lars Carlsson
    Abstract:

    : Valid and predictive models for classifying Ames mutagenicity have been developed using Conformal Prediction. The models are Random Forest models using signature molecular descriptors. The investigation indicates, on excluding not-strongly mutagenic compounds (class B), that the validity for mutagenic compounds is increased for the Predictions based on both public and the Division of Genetics and Mutagenesis, National Institute of Health Sciences of Japan (DGM/NIHS) data while less so when using only the latter data source. The former models only result in valid Predictions for the majority, non-mutagenic, class whereas the latter models are valid for both classes, i.e. mutagenic and non-mutagenic compounds. These results demonstrate the importance of data consistency manifested through the superior predictive quality and validity of the models based only on DGM/NIHS generated data compared to a combination of this data with public data sources.

  • Predicting Ames Mutagenicity Using Conformal Prediction in the Ames/QSAR International Challenge Project.
    Mutagenesis, 2018
    Co-Authors: Ulf Norinder, Ernst Ahlberg, Lars Carlsson
    Abstract:

    : Valid and predictive models for classifying Ames mutagenicity have been developed using Conformal Prediction. The models are Random Forest models using signature molecular descriptors. The investigation indicates, on excluding not-strongly mutagenic compounds (class B), that the validity for mutagenic compounds is increased for the Predictions based on both public and the Division of Genetics and Mutagenesis, National Institute of Health Sciences of Japan (DGM/NIHS) data while less so when using only the latter data source. The former models only result in valid Predictions for the majority, non-mutagenic, class whereas the latter models are valid for both classes, i.e. mutagenic and non-mutagenic compounds. These results demonstrate the importance of data consistency manifested through the superior predictive quality and validity of the models based only on DGM/NIHS generated data compared to a combination of this data with public data sources.

  • Aggregating Predictions on Multiple Non-disclosed Datasets using Conformal Prediction.
    arXiv: Machine Learning, 2018
    Co-Authors: Ola Spjuth, Lars Carlsson, Niharika Gauraha
    Abstract:

    Conformal Prediction is a machine learning methodology that produces valid Prediction regions under mild conditions. In this paper, we explore the application of making Predictions over multiple data sources of different sizes without disclosing data between the sources. We propose that each data source applies a transductive Conformal predictor independently using the local data, and that the individual Predictions are then aggregated to form a combined Prediction region. We demonstrate the method on several data sets, and show that the proposed method produces conservatively valid Predictions and reduces the variance in the aggregated Predictions. We also study the effect that the number of data sources and size of each source has on aggregated Predictions, as compared with equally sized sources and pooled data.

  • Conformal Prediction in learning under privileged information paradigm with applications in drug discovery
    arXiv: Machine Learning, 2018
    Co-Authors: Niharika Gauraha, Lars Carlsson, Ola Spjuth
    Abstract:

    This paper explores Conformal Prediction in the learning under privileged information (LUPI) paradigm. We use the SVM+ realization of LUPI in an inductive Conformal predictor, and apply it to the MNIST benchmark dataset and three datasets in drug discovery. The results show that using privileged information produces valid models and improves efficiency compared to standard SVM, however the improvement varies between the tested datasets and is not substantial in the drug discovery applications. More importantly, using SVM+ in a Conformal Prediction framework enables valid Prediction intervals at specified significance levels.

  • Current application of Conformal Prediction in drug discovery
    Annals of Mathematics and Artificial Intelligence, 2017
    Co-Authors: Ernst Ahlberg, Oscar Hammar, Claus Bendtsen, Lars Carlsson
    Abstract:

    We present two applications of Conformal Prediction relevant to drug discovery. The first application is around interpretation of Predictions and the second one around the selection of compounds to progress in a drug discovery project setting.

Alexander Gammerman - One of the best experts on this subject based on the ideXlab platform.

  • ICCSW - Conformal Prediction under Hypergraphical Models
    2020
    Co-Authors: Valentina Fedorova, Ilia Nouretdinov, Alexander Gammerman, Vladimir Vovk
    Abstract:

    Conformal predictors are usually defined and studied under the exchangeability assumption. However, their definition can be extended to a wide class of statistical models, called online compression models, while retaining their property of automatic validity. This paper is devoted to Conformal Prediction under hypergraphical models that are more specific than the exchangeability model. We define conformity measures for such hypergraphical models and study the corresponding Conformal predictors empirically on benchmark LED data sets. Our experiments show that they are more efficient than Conformal predictors that use only the exchangeability assumption.

  • criteria of efficiency for Conformal Prediction
    COPA 2016 Proceedings of the 5th International Symposium on Conformal and Probabilistic Prediction with Applications - Volume 9653, 2016
    Co-Authors: Vladimir Vovk, Ilia Nouretdinov, Valentina Fedorova, Alexander Gammerman
    Abstract:

    We study optimal conformity measures for various criteria of efficiency in an idealised setting. This leads to an important class of criteria of efficiency that we call probabilistic; it turns out that the most standard criteria of efficiency used in literature on Conformal Prediction are not probabilistic.

  • COPA - Criteria of Efficiency for Conformal Prediction
    Lecture Notes in Computer Science, 2016
    Co-Authors: Vladimir Vovk, Ilia Nouretdinov, Valentina Fedorova, Alexander Gammerman
    Abstract:

    We study optimal conformity measures for various criteria of efficiency in an idealised setting. This leads to an important class of criteria of efficiency that we call probabilistic; it turns out that the most standard criteria of efficiency used in literature on Conformal Prediction are not probabilistic.

  • criteria of efficiency for Conformal Prediction
    arXiv: Learning, 2016
    Co-Authors: Vladimir Vovk, Ilia Nouretdinov, Valentina Fedorova, Ivan Petej, Alexander Gammerman
    Abstract:

    We study optimal conformity measures for various criteria of efficiency of classification in an idealised setting. This leads to an important class of criteria of efficiency that we call probabilistic; it turns out that the most standard criteria of efficiency used in literature on Conformal Prediction are not probabilistic unless the problem of classification is binary. We consider both unconditional and label-conditional Conformal Prediction.

  • Anomaly Detection of Trajectories with Kernel Density Estimation by Conformal Prediction
    2014
    Co-Authors: James Smith, Charles Offer, Ilia Nouretdinov, Rachel Craddock, Alexander Gammerman
    Abstract:

    This paper describes Conformal Prediction techniques for detecting anomalous trajectories in the maritime domain. The data used in experiments were obtained from Automatic Identification System (AIS) broadcasts – a system for tracking vessel locations. A dimensionality reduction package is used and a kernel density estimation function as a non-conformity measure has been applied to detect anomalies. We propose average p-value as an efficiency criteria for Conformal anomaly detection. A comparison with a k-nearest neighbours non-conformity measure is presented and the results are discussed.