The Experts below are selected from a list of 101781 Experts worldwide ranked by ideXlab platform

Miron B. Kursa - One of the best experts on this subject based on the ideXlab platform.

  • robustness of Random Forest based gene selection methods
    BMC Bioinformatics, 2014
    Co-Authors: Miron B. Kursa
    Abstract:

    Background Gene selection is an important part of microarray data analysis because it provides information that can lead to a better mechanistic understanding of an investigated phenomenon. At the same time, gene selection is very difficult because of the noisy nature of microarray data. As a consequence, gene selection is often performed with machine learning methods. The Random Forest method is particularly well suited for this purpose. In this work, four state-of-the-art Random Forest-based feature selection methods were compared in a gene selection context. The analysis focused on the stability of selection because, although it is necessary for determining the significance of results, it is often ignored in similar studies.

  • Robustness of Random Forest-based gene selection methods
    BMC Bioinformatics, 2014
    Co-Authors: Miron B. Kursa
    Abstract:

    Gene selection is an important part of microarray data analysis because it provides information that can lead to a better mechanistic understanding of an investigated phenomenon. At the same time, gene selection is very difficult because of the noisy nature of microarray data. As a consequence, gene selection is often performed with machine learning methods. The Random Forest method is particularly well suited for this purpose. In this work, four state-of-the-art Random Forest-based feature selection methods were compared in a gene selection context. The analysis focused on the stability of selection because, although it is necessary for determining the significance of results, it is often ignored in similar studies. The comparison of post-selection accuracy in the validation of Random Forest classifiers revealed that all investigated methods were equivalent in this context. However, the methods substantially differed with respect to the number of selected genes and the stability of selection. Of the analysed methods, the Boruta algorithm predicted the most genes as potentially important. The post-selection classifier error rate, which is a frequently used measure, was found to be a potentially deceptive measure of gene selection quality. When the number of consistently selected genes was considered, the Boruta algorithm was clearly the best. Although it was also the most computationally intensive method, the Boruta algorithm's computational demands could be reduced to levels comparable to those of other algorithms by replacing the Random Forest importance with a comparable measure from Random Ferns (a similar but simplified classifier). Despite their design assumptions, the minimal optimal selection methods, were found to select a high fraction of false positives.

  • ISMIS - Musical Instruments in Random Forest
    Lecture Notes in Computer Science, 2009
    Co-Authors: Miron B. Kursa, Witold R. Rudnicki, Alicja Wieczorkowska, Elżbieta Kubera, Agnieszka Kubik-komar
    Abstract:

    This paper describes automatic classification of predominant musical instrument in sound mixes, using Random Forests as classifiers. The description of sound parameterization applied and methodology of Random Forest classification are given in the paper. Additionally, the significance of sound parameters used as conditional attributes is investigated. The results show that almost all sound attributes are informative, and Random Forest technique yields much higher classification results than support vector machines, used in previous research on these data.

Frederick Livingston - One of the best experts on this subject based on the ideXlab platform.

Sara Alvarez De Andrés - One of the best experts on this subject based on the ideXlab platform.

  • Gene selection and classification of microarray data using Random Forest
    BMC Bioinformatics, 2006
    Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De Andrés
    Abstract:

    Background Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. Results We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. Conclusion Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.

  • Gene selection and classification of microarray data using Random Forest
    BMC Bioinformatics, 2006
    Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De Andrés
    Abstract:

    BACKGROUND: Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. RESULTS: We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. CONCLUSION: Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.

Ramón Díaz-uriarte - One of the best experts on this subject based on the ideXlab platform.

  • Gene selection and classification of microarray data using Random Forest
    BMC Bioinformatics, 2006
    Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De Andrés
    Abstract:

    Background Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. Results We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. Conclusion Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.

  • Gene selection and classification of microarray data using Random Forest
    BMC Bioinformatics, 2006
    Co-Authors: Ramón Díaz-uriarte, Sara Alvarez De Andrés
    Abstract:

    BACKGROUND: Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, Random Forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of Random Forest with microarray data and its possible use for gene selection. RESULTS: We investigate the use of Random Forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on Random Forest. Using simulated and nine microarray data sets we show that Random Forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. CONCLUSION: Because of its performance and features, Random Forest and gene selection using Random Forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.

Abdulhamit Subasi - One of the best experts on this subject based on the ideXlab platform.

  • congestive heart failure detection using Random Forest classifier
    Computer Methods and Programs in Biomedicine, 2016
    Co-Authors: Zerina Masetic, Abdulhamit Subasi
    Abstract:

    Heartbeat classification is substantial for diagnosing heart failure.Machine learning methods classify normal and congestive heart failure (CHF).The Random Forest method gives 100% classification accuracy in detecting CHF. Background and objectivesAutomatic electrocardiogram (ECG) heartbeat classification is substantial for diagnosing heart failure. The aim of this paper is to evaluate the effect of machine learning methods in creating the model which classifies normal and congestive heart failure (CHF) on the long-term ECG time series. MethodsThe study was performed in two phases: feature extraction and classification phase. In feature extraction phase, autoregressive (AR) Burg method is applied for extracting features. In classification phase, five different classifiers are examined namely, C4.5 decision tree, k-nearest neighbor, support vector machine, artificial neural networks and Random Forest classifier. The ECG signals were acquired from BIDMC Congestive Heart Failure and PTB Diagnostic ECG databases and classified by applying various experiments. ResultsThe experimental results are evaluated in several statistical measures (sensitivity, specificity, accuracy, F-measure and ROC curve) and showed that the Random Forest method gives 100% classification accuracy. ConclusionsImpressive performance of Random Forest method proves that it plays significant role in detecting congestive heart failure (CHF) and can be valuable in expressing knowledge useful in medicine.

  • Congestive heart failure detection using Random Forest classifier
    Computer Methods and Programs in Biomedicine, 2016
    Co-Authors: Zerina Masetic, Abdulhamit Subasi
    Abstract:

    Background and objectives: Automatic electrocardiogram (ECG) heartbeat classification is substantial for diagnosing heart failure. The aim of this paper is to evaluate the effect of machine learning methods in creating the model which classifies normal and congestive heart failure (CHF) on the long-term ECG time series. Methods: The study was performed in two phases: feature extraction and classification phase. In feature extraction phase, autoregressive (AR) Burg method is applied for extracting features. In classification phase, five different classifiers are examined namely, C4.5 decision tree, k-nearest neighbor, support vector machine, artificial neural networks and Random Forest classifier. The ECG signals were acquired from BIDMC Congestive Heart Failure and PTB Diagnostic ECG databases and classified by applying various experiments. Results: The experimental results are evaluated in several statistical measures (sensitivity, specificity, accuracy, F-measure and ROC curve) and showed that the Random Forest method gives 100% classification accuracy. Conclusions: Impressive performance of Random Forest method proves that it plays significant role in detecting congestive heart failure (CHF) and can be valuable in expressing knowledge useful in medicine.