The Experts below are selected from a list of 17787 Experts worldwide ranked by ideXlab platform
Jocelyn Chanussot - One of the best experts on this subject based on the ideXlab platform.
-
hyperspectral remote sensing image classification Based on rotation forest
IEEE Geoscience and Remote Sensing Letters, 2014Co-Authors: Peijun Du, Xiyan He, Jocelyn ChanussotAbstract:In this letter, an ensemble learning approach, Rotation Forest, has been applied to hyperspectral remote sensing image classification for the first time. The framework of Rotation Forest is to project the original data into a new feature space using transformation methods for each Base Classifier (decision tree), then the Base Classifier can train in different new spaces for the purpose of encouraging both individual accuracy and diversity within the ensemble simultaneously. Principal component analysis (PCA), maximum noise fraction, independent component analysis, and local Fisher discriminant analysis are introduced as feature transformation algorithms in the original Rotation Forest. The performance of Rotation Forest was evaluated Based on several criteria: different data sets, sensitivity to the number of training samples, ensemble size and the number of features in a subset. Experimental results revealed that Rotation Forest, especially with PCA transformation, could produce more accurate results than bagging, AdaBoost, and Random Forest. They indicate that Rotation Forests are promising approaches for generating Classifier ensemble of hyperspectral remote sensing.
Peijun Du - One of the best experts on this subject based on the ideXlab platform.
-
hyperspectral remote sensing image classification Based on rotation forest
IEEE Geoscience and Remote Sensing Letters, 2014Co-Authors: Peijun Du, Xiyan He, Jocelyn ChanussotAbstract:In this letter, an ensemble learning approach, Rotation Forest, has been applied to hyperspectral remote sensing image classification for the first time. The framework of Rotation Forest is to project the original data into a new feature space using transformation methods for each Base Classifier (decision tree), then the Base Classifier can train in different new spaces for the purpose of encouraging both individual accuracy and diversity within the ensemble simultaneously. Principal component analysis (PCA), maximum noise fraction, independent component analysis, and local Fisher discriminant analysis are introduced as feature transformation algorithms in the original Rotation Forest. The performance of Rotation Forest was evaluated Based on several criteria: different data sets, sensitivity to the number of training samples, ensemble size and the number of features in a subset. Experimental results revealed that Rotation Forest, especially with PCA transformation, could produce more accurate results than bagging, AdaBoost, and Random Forest. They indicate that Rotation Forests are promising approaches for generating Classifier ensemble of hyperspectral remote sensing.
Brijesh Verma - One of the best experts on this subject based on the ideXlab platform.
-
A Novel Diversity Measure and Classifier Selection Approach for Generating Ensemble Classifiers
IEEE Access, 2019Co-Authors: Muhammad Zohaib Jan, Brijesh VermaAbstract:Accuracy and diversity are considered to be the two deriving factors when it comes to generating an ensemble Classifier. Focusing only on accuracy causes the ensemble Classifier to suffer from “diminishing returns” and the ensemble accuracy tends to plateau; whereas focusing only on diversity causes the ensemble Classifier to suffer in accuracy. Therefore, a balance must be maintained between the two for the ensemble Classifier to achieve high classification accuracy. In this paper, we propose a novel diversity measure known as Misclassification Diversity (MD) and an Incremental Layered Classifier Selection (ILCS) approach to generate an ensemble Classifier. The proposed approach ILCS-MD generates an ensemble Classifier by incrementally selecting Classifiers from the Base Classifier pool Based on increasing accuracy and diversity. The benefits are in two folds 1) the generated ensemble Classifier contains only those Classifiers from the pool which can either maximize accuracy whilst maintaining or increasing the diversity, and 2) the generated ensemble Classifier selects only a few Classifiers from the Base Classifier pool thus reducing ensemble component size as well. The proposed approach is evaluated on 55 benchmark datasets taken from UCI and KEEL dataset repositories. The results are compared with five existing pairwise diversity measures, and existing state of the art ensemble Classifier approaches. A significance test is also conducted to verify the significance of the results.
-
ensemble Classifier optimization by reducing input features and Base Classifiers
Congress on Evolutionary Computation, 2019Co-Authors: Zohaib Jan, Brijesh VermaAbstract:Ensemble Classifier approaches either exploit the input feature space also known as the dataset attributes, or exploit the data sample space, for example Random Forest (RaF) exploits input features whereas Bagging exploits data sample space. Very few ensemble Classifier approaches exist that exploit the both. In this paper we propose an ensemble Classifier approach that first reduces input feature space by selecting only the significant features from input data that can maximize the classification performance of the ensemble and then optimize the Base Classifier pool by incorporating an evolutionary algorithm. The proposed approach is evaluated on benchmark datasets from UCI repository. The results are compared with single Classifier approaches and existing state of the art ensemble Classifier approaches.
Jason Weston - One of the best experts on this subject based on the ideXlab platform.
-
combining labeled and unlabeled data with word class distribution learning
Conference on Information and Knowledge Management, 2009Co-Authors: Ronan Collobert, Pavel P Kuksa, Koray Kavukcuoglu, Jason WestonAbstract:We describe a novel simple and highly scalable semi-supervised method called Word-Class Distribution Learning (WCDL), and apply it task of information extraction (IE) by utilizing unlabeled sentences to improve supervised classification methods. WCDL iteratively builds class label distributions for each word in the dictionary by averaging predicted labels over all cases in the unlabeled corpus, and re-training a Base Classifier adding these distributions as word features. In contrast, traditional self-training or co-training methods self-labeled examples (rather than features) which can degrade performance due to incestuous learning bias. WCDL exhibits robust behavior, and has no difficult parameters to tune. We applied our method on German and English name entity recognition (NER) tasks. WCDL shows improvements over self-training, multi-task semi-supervision or supervision alone, in particular yielding a state-of-the art 75.72 F1 score on the German NER task.
-
CIKM - Combining labeled and unlabeled data with word-class distribution learning
Proceeding of the 18th ACM conference on Information and knowledge management - CIKM '09, 2009Co-Authors: Ronan Collobert, Pavel P Kuksa, Koray Kavukcuoglu, Jason WestonAbstract:We describe a novel simple and highly scalable semi-supervised method called Word-Class Distribution Learning (WCDL), and apply it task of information extraction (IE) by utilizing unlabeled sentences to improve supervised classification methods. WCDL iteratively builds class label distributions for each word in the dictionary by averaging predicted labels over all cases in the unlabeled corpus, and re-training a Base Classifier adding these distributions as word features. In contrast, traditional self-training or co-training methods self-labeled examples (rather than features) which can degrade performance due to incestuous learning bias. WCDL exhibits robust behavior, and has no difficult parameters to tune. We applied our method on German and English name entity recognition (NER) tasks. WCDL shows improvements over self-training, multi-task semi-supervision or supervision alone, in particular yielding a state-of-the art 75.72 F1 score on the German NER task.
Joaquín Abellán - One of the best experts on this subject based on the ideXlab platform.
-
Ensemble of Classifier chains and Credal C4.5 for solving multi-label classification
Progress in Artificial Intelligence, 2019Co-Authors: Serafín Moral-garcía, Carlos Javier Mantas, Javier G. Castellano, Joaquín AbellánAbstract:In this work, we have considered the ensemble of Classifier chains (ECC) algorithm in order to solve the multi-label classification (MLC) task. It starts from binary relevance algorithm (BR), a simple and direct approach to MLC that has been shown to provide good results in practice. Nevertheless, unlike BR, ECC aims to exploit the correlations between labels. ECC uses an algorithm of traditional supervised classification in order to approach the binary problems. Within this field, Credal C4.5 (CC4.5) is a new version of the well-known C4.5 algorithm that uses imprecise probabilities in order to estimate the probability distribution of the class variable. This new version of C4.5 algorithm has been shown to provide better performance when noisy datasets are classified. In MLC, the intrinsic noise might be higher than in traditional supervised classification. The reason is very simple: in MLC, there are multiple labels, whereas in traditional classification there is just a class variable. Thus, there is more probability of error for an instance. For the previous reasons, the performance of ECC with CC4.5 as Base Classifier is studied in this work. We have carried out an extensive experimental analysis with several multi-label datasets, different noise levels and a large number of evaluation metrics for MLC. This experimental study has shown that, generally, ECC has better performance with CC4.5 as Base Classifier than using C4.5. The higher is the label noise level introduced in the data, the more significative is this improvement. Therefore, it is probably suitable to use imprecise probabilities in Decision Trees within MLC.
-
Ensemble of Classifier chains and Credal C4.5 for solving multi-label classification
Progress in Artificial Intelligence, 2019Co-Authors: Serafín Moral-garcía, Carlos Javier Mantas, Javier G. Castellano, Joaquín AbellánAbstract:In this work, we have considered the ensemble of Classifier chains (ECC) algorithm in order to solve the multi-label classification (MLC) task. It starts from binary relevance algorithm (BR), a simple and direct approach to MLC that has been shown to provide good results in practice. Nevertheless, unlike BR, ECC aims to exploit the correlations between labels. ECC uses an algorithm of traditional supervised classification in order to approach the binary problems. Within this field, Credal C4.5 (CC4.5) is a new version of the well-known C4.5 algorithm that uses imprecise probabilities in order to estimate the probability distribution of the class variable. This new version of C4.5 algorithm has been shown to provide better performance when noisy datasets are classified. In MLC, the intrinsic noise might be higher than in traditional supervised classification. The reason is very simple: in MLC, there are multiple labels, whereas in traditional classification there is just a class variable. Thus, there is more probability of error for an instance. For the previous reasons, the performance of ECC with CC4.5 as Base Classifier is studied in this work. We have carried out an extensive experimental analysis with several multi-label datasets, different noise levels and a large number of evaluation metrics for MLC. This experimental study has shown that, generally, ECC has better performance with CC4.5 as Base Classifier than using C4.5. The higher is the label noise level introduced in the data, the more significative is this improvement. Therefore, it is probably suitable to use imprecise probabilities in Decision Trees within MLC.
-
A comparative study on Base Classifiers in ensemble methods for credit scoring
Expert Systems with Applications, 2017Co-Authors: Joaquín Abellán, Javier G. CastellanoAbstract:Abstract In the last years, the application of artificial intelligence methods on credit risk assessment has meant an improvement over classic methods. Small improvements in the systems about credit scoring and bankruptcy prediction can suppose great profits. Then, any improvement represents a high interest to banks and financial institutions. Recent works show that ensembles of Classifiers achieve the better results for this kind of tasks. In this paper, it is extended a previous work about the selection of the best Base Classifier used in ensembles on credit data sets. It is shown that a very simple Base Classifier, Based on imprecise probabilities and uncertainty measures, attains a better trade-off among some aspects of interest for this type of studies such as accuracy and area under ROC curve (AUC). The AUC measure can be considered as a more appropriate measure in this grounds, where the different type of errors have different costs or consequences. The results shown here present to this simple Classifier as an interesting choice to be used as Base Classifier in ensembles for credit scoring and bankruptcy prediction, proving that not only the individual performance of a Classifier is the key point to be selected for an ensemble scheme.