The Experts below are selected from a list of 243 Experts worldwide ranked by ideXlab platform
Louis Wehenkel - One of the best experts on this subject based on the ideXlab platform.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a new tree-based ensemble method for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency. A bias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a newtree-based ensemblemethod for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency.Abias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced. Keywords
Pierre Geurts - One of the best experts on this subject based on the ideXlab platform.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a new tree-based ensemble method for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency. A bias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a newtree-based ensemblemethod for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency.Abias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced. Keywords
Damien Ernst - One of the best experts on this subject based on the ideXlab platform.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a new tree-based ensemble method for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency. A bias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced.
-
Extremely randomized trees
Machine Learning, 2006Co-Authors: Pierre Geurts, Damien Ernst, Louis WehenkelAbstract:This paper proposes a newtree-based ensemblemethod for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the Learning Sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency.Abias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced. Keywords
Maria Samsonova - One of the best experts on this subject based on the ideXlab platform.
-
A regression system for estimation of errors introduced by confocal imaging into gene expression data in situ
BMC Bioinformatics, 2011Co-Authors: Ekaterina Myasnikova, Svetlana Surkova, Grigory Stein, Andrei Pisarev, Maria SamsonovaAbstract:Background Accuracy of the data extracted from two-dimensional confocal images is limited due to experimental errors that arise in course of confocal scanning. The common way to reduce the noise in images is sequential scanning of the same specimen several times with the subsequent averaging of multiple frames. Attempts to increase the dynamical range of an image by setting too high values of microscope PMT parameters may cause clipping of single frames and introduce errors into the data extracted from the averaged images. For the estimation and correction of this kind of errors a method based on censoring technique (Myasnikova et al., 2009) is used. However, the method requires the availability of all the confocal scans along with the averaged image, which is normally not provided by the standard scanning procedure. Results To predict error size in the data extracted from the averaged image we developed a regression system. The system is trained on the Learning Sample composed of images obtained from three different microscopes at different combinations of PMT parameters, and for each image all the scans are saved. The system demonstrates high prediction accuracy and was applied for correction of errors in the data on segmentation gene expression in Drosophila blastoderm stored in the FlyEx database ( http://urchin.spbcas.ru/flyex/ , http://flyex.uchicago.edu/flyex/ ). The prediction method is realized as a software tool CorrectPattern freely available at http://urchin.spbcas.ru/asp/2011/emm/ . Conclusions We created a regression system and software to predict the magnitude of errors in the data obtained from a confocal image based on information about microscope parameters used for the image acquisition. An important advantage of the developed prediction system is the possibility to accurately correct the errors in data obtained from strongly clipped images, thereby allowing to obtain images of the higher dynamical range and thus to extract more detailed quantitative information from them.
Eric P. Xing - One of the best experts on this subject based on the ideXlab platform.
-
NeurIPS - Learning Sample-Specific Models with Low-Rank Personalized Regression
2019Co-Authors: Benjamin J. Lengerich, Bryon Aragam, Eric P. XingAbstract:Modern applications of machine Learning (ML) deal with increasingly heterogeneous datasets comprised of data collected from overlapping latent subpopulations. As a result, traditional models trained over large datasets may fail to recognize highly predictive localized effects in favour of weakly predictive global patterns. This is a problem because localized effects are critical to developing individualized policies and treatment plans in applications ranging from precision medicine to advertising. To address this challenge, we propose to estimate Sample-specific models that tailor inference and prediction at the individual level. In contrast to classical ML models that estimate a single, complex model (or only a few complex models), our approach produces a model personalized to each Sample. These Sample-specific models can be studied to understand subgroup dynamics that go beyond coarse-grained class labels. Crucially, our approach does not assume that relationships between Samples (e.g. a similarity network) are known a priori. Instead, we use unmodeled covariates to learn a latent distance metric over the Samples. We apply this approach to financial, biomedical, and electoral data as well as simulated data and show that Sample-specific models provide fine-grained interpretations of complicated phenomena without sacrificing predictive accuracy compared to state-of-the-art models such as deep neural networks.
-
Learning Sample-Specific Models with Low-Rank Personalized Regression
arXiv: Machine Learning, 2019Co-Authors: Benjamin J. Lengerich, Bryon Aragam, Eric P. XingAbstract:Modern applications of machine Learning (ML) deal with increasingly heterogeneous datasets comprised of data collected from overlapping latent subpopulations. As a result, traditional models trained over large datasets may fail to recognize highly predictive localized effects in favour of weakly predictive global patterns. This is a problem because localized effects are critical to developing individualized policies and treatment plans in applications ranging from precision medicine to advertising. To address this challenge, we propose to estimate Sample-specific models that tailor inference and prediction at the individual level. In contrast to classical ML models that estimate a single, complex model (or only a few complex models), our approach produces a model personalized to each Sample. These Sample-specific models can be studied to understand subgroup dynamics that go beyond coarse-grained class labels. Crucially, our approach does not assume that relationships between Samples (e.g. a similarity network) are known a priori. Instead, we use unmodeled covariates to learn a latent distance metric over the Samples. We apply this approach to financial, biomedical, and electoral data as well as simulated data and show that Sample-specific models provide fine-grained interpretations of complicated phenomena without sacrificing predictive accuracy compared to state-of-the-art models such as deep neural networks.