The Experts below are selected from a list of 101781 Experts worldwide ranked by ideXlab platform
Miron B. Kursa - One of the best experts on this subject based on the ideXlab platform.
-
robustness of Random Forest based gene selection methods
BMC Bioinformatics, 2014Co-Authors: Miron B. KursaAbstract:Background Gene selection is an important part of microarray data analysis because it provides information that can lead to a better mechanistic understanding of an investigated phenomenon. At the same time, gene selection is very difficult because of the noisy nature of microarray data. As a consequence, gene selection is often performed with machine learning methods. The Random Forest method is particularly well suited for this purpose. In this work, four state-of-the-art Random Forest-based feature selection methods were compared in a gene selection context. The analysis focused on the stability of selection because, although it is necessary for determining the significance of results, it is often ignored in similar studies.
-
Robustness of Random Forest-based gene selection methods
BMC Bioinformatics, 2014Co-Authors: Miron B. KursaAbstract:Gene selection is an important part of microarray data analysis because it provides information that can lead to a better mechanistic understanding of an investigated phenomenon. At the same time, gene selection is very difficult because of the noisy nature of microarray data. As a consequence, gene selection is often performed with machine learning methods. The Random Forest method is particularly well suited for this purpose. In this work, four state-of-the-art Random Forest-based feature selection methods were compared in a gene selection context. The analysis focused on the stability of selection because, although it is necessary for determining the significance of results, it is often ignored in similar studies. The comparison of post-selection accuracy in the validation of Random Forest classifiers revealed that all investigated methods were equivalent in this context. However, the methods substantially differed with respect to the number of selected genes and the stability of selection. Of the analysed methods, the Boruta algorithm predicted the most genes as potentially important. The post-selection classifier error rate, which is a frequently used measure, was found to be a potentially deceptive measure of gene selection quality. When the number of consistently selected genes was considered, the Boruta algorithm was clearly the best. Although it was also the most computationally intensive method, the Boruta algorithm's computational demands could be reduced to levels comparable to those of other algorithms by replacing the Random Forest importance with a comparable measure from Random Ferns (a similar but simplified classifier). Despite their design assumptions, the minimal optimal selection methods, were found to select a high fraction of false positives.
-
ISMIS - Musical Instruments in Random Forest
Lecture Notes in Computer Science, 2009Co-Authors: Miron B. Kursa, Witold R. Rudnicki, Alicja Wieczorkowska, Elżbieta Kubera, Agnieszka Kubik-komarAbstract:This paper describes automatic classification of predominant musical instrument in sound mixes, using Random Forests as classifiers. The description of sound parameterization applied and methodology of Random Forest classification are given in the paper. Additionally, the significance of sound parameters used as conditional attributes is investigated. The results show that almost all sound attributes are informative, and Random Forest technique yields much higher classification results than support vector machines, used in previous research on these data.
Frederick Livingston - One of the best experts on this subject based on the ideXlab platform.
-
Implementation of Breiman’s Random Forest Machine Learning Algorithm
Machine Learning Journal Paper, 2005Co-Authors: Frederick LivingstonAbstract:This research provides tools for exploring Breiman’s Random Forest algorithm. This paper will focus on the development, the verification, and the significance of variable importance.
-
Implementation of Breiman Random Forest Machine Learning Algorithm
Machine Learning Journal Paper, 2005Co-Authors: Frederick LivingstonAbstract:This research provides tools for exploring Breiman Random Forest algorithm. This paper will focus on the development, the verification, and the significance of variable importance.
Sara Alvarez De Andrés - One of the best experts on this subject based on the ideXlab platform.
-
Gene selection and classification of microarray data using Random Forest
BMC Bioinformatics, 2006Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De AndrésAbstract:Background Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. Results We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. Conclusion Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.
-
Gene selection and classification of microarray data using Random Forest
BMC Bioinformatics, 2006Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De AndrésAbstract:BACKGROUND: Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. RESULTS: We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. CONCLUSION: Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.
Ramón Díaz-uriarte - One of the best experts on this subject based on the ideXlab platform.
-
Gene selection and classification of microarray data using Random Forest
BMC Bioinformatics, 2006Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De AndrésAbstract:Background Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. Results We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. Conclusion Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.
-
Gene selection and classification of microarray data using Random Forest
BMC Bioinformatics, 2006Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De AndrésAbstract:BACKGROUND: Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. RESULTS: We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. CONCLUSION: Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.
Abdulhamit Subasi - One of the best experts on this subject based on the ideXlab platform.
-
congestive heart failure detection using Random Forest classifier
Computer Methods and Programs in Biomedicine, 2016Co-Authors: Zerina Masetic, Abdulhamit SubasiAbstract:Heartbeat classification is substantial for diagnosing heart failure.Machine learning methods classify normal and congestive heart failure (CHF).The Random Forest method gives 100% classification accuracy in detecting CHF. Background and objectivesAutomatic electrocardiogram (ECG) heartbeat classification is substantial for diagnosing heart failure. The aim of this paper is to evaluate the effect of machine learning methods in creating the model which classifies normal and congestive heart failure (CHF) on the long-term ECG time series. MethodsThe study was performed in two phases: feature extraction and classification phase. In feature extraction phase, autoregressive (AR) Burg method is applied for extracting features. In classification phase, five different classifiers are examined namely, C4.5 decision tree, k-nearest neighbor, support vector machine, artificial neural networks and Random Forest classifier. The ECG signals were acquired from BIDMC Congestive Heart Failure and PTB Diagnostic ECG databases and classified by applying various experiments. ResultsThe experimental results are evaluated in several statistical measures (sensitivity, specificity, accuracy, F-measure and ROC curve) and showed that the Random Forest method gives 100% classification accuracy. ConclusionsImpressive performance of Random Forest method proves that it plays significant role in detecting congestive heart failure (CHF) and can be valuable in expressing knowledge useful in medicine.
-
Congestive heart failure detection using Random Forest classifier
Computer Methods and Programs in Biomedicine, 2016Co-Authors: Zerina Masetic, Abdulhamit SubasiAbstract:Background and objectives: Automatic electrocardiogram (ECG) heartbeat classification is substantial for diagnosing heart failure. The aim of this paper is to evaluate the effect of machine learning methods in creating the model which classifies normal and congestive heart failure (CHF) on the long-term ECG time series. Methods: The study was performed in two phases: feature extraction and classification phase. In feature extraction phase, autoregressive (AR) Burg method is applied for extracting features. In classification phase, five different classifiers are examined namely, C4.5 decision tree, k-nearest neighbor, support vector machine, artificial neural networks and Random Forest classifier. The ECG signals were acquired from BIDMC Congestive Heart Failure and PTB Diagnostic ECG databases and classified by applying various experiments. Results: The experimental results are evaluated in several statistical measures (sensitivity, specificity, accuracy, F-measure and ROC curve) and showed that the Random Forest method gives 100% classification accuracy. Conclusions: Impressive performance of Random Forest method proves that it plays significant role in detecting congestive heart failure (CHF) and can be valuable in expressing knowledge useful in medicine.