The Experts below are selected from a list of 3423 Experts worldwide ranked by ideXlab platform
Jinxing Che - One of the best experts on this subject based on the ideXlab platform.
-
maximum relevance minimum common Redundancy feature selection for nonlinear data
Information Sciences, 2017Co-Authors: Jinxing Che, Youlong Yang, Xuying Bai, Shenghu Zhang, Chengzhi DengAbstract:In recent years, feature selection based on relevance Redundancy trade-off criteria has become a very promising and popular approach in the field of machine learning. However, the existing algorithmic frameworks of mutual information feature selection have certain limitations for the common feature selection problems in practice. To overcome these limitations, the idea of a new framework is developed by introducing a novel maximum relevance and minimum common Redundancy Criterion and a minimax nonlinear optimization approach. In particular, a novel mutual information feature selection method based on the normalization of the maximum relevance and minimum common Redundancy (N-MRMCR-MI) is presented, which produces a normalized value in the range [0, 1] and results in a regression problem. We perform extensive experimental comparisons over numerous state-of-art algorithms using different forecasts (Bayesian Additive Regression tree, treed Gaussian process, k-NN, and SVM) and different data sets (two simulated and five real datasets). The results show that the proposed algorithm outperforms the others in terms of feature selection and forecasting accuracy.
Chengzhi Deng - One of the best experts on this subject based on the ideXlab platform.
-
maximum relevance minimum common Redundancy feature selection for nonlinear data
Information Sciences, 2017Co-Authors: Jinxing Che, Youlong Yang, Xuying Bai, Shenghu Zhang, Chengzhi DengAbstract:In recent years, feature selection based on relevance Redundancy trade-off criteria has become a very promising and popular approach in the field of machine learning. However, the existing algorithmic frameworks of mutual information feature selection have certain limitations for the common feature selection problems in practice. To overcome these limitations, the idea of a new framework is developed by introducing a novel maximum relevance and minimum common Redundancy Criterion and a minimax nonlinear optimization approach. In particular, a novel mutual information feature selection method based on the normalization of the maximum relevance and minimum common Redundancy (N-MRMCR-MI) is presented, which produces a normalized value in the range [0, 1] and results in a regression problem. We perform extensive experimental comparisons over numerous state-of-art algorithms using different forecasts (Bayesian Additive Regression tree, treed Gaussian process, k-NN, and SVM) and different data sets (two simulated and five real datasets). The results show that the proposed algorithm outperforms the others in terms of feature selection and forecasting accuracy.
Ujjwal Maulik - One of the best experts on this subject based on the ideXlab platform.
-
identifying epigenetic biomarkers using maximal relevance and minimal Redundancy based feature selection for multi omics data
IEEE Transactions on Nanobioscience, 2017Co-Authors: Saurav Mallik, Tapas Bhadra, Ujjwal MaulikAbstract:Epigenetic Biomarker discovery is an important task in bioinformatics. In this article, we develop a new framework of identifying statistically significant epigenetic biomarkers using maximal-relevance and minimal-Redundancy Criterion based feature (gene) selection for multi-omics dataset. Firstly, we determine the genes that have both expression as well as methylation values, and follow normal distribution. Similarly, we identify the genes which consist of both expression and methylation values, but do not follow normal distribution. For each case, we utilize a gene-selection method that provides maximal-relevant, but variable-weighted minimum-redundant genes as top ranked genes. For statistical validation, we apply t-test on both the expression and methylation data consisting of only the normally distributed top ranked genes to determine how many of them are both differentially expressed andmethylated. Similarly, we utilize Limma package for performing non-parametric Empirical Bayes test on both expression and methylation data comprising only the non-normally distributed top ranked genes to identify how many of them are both differentially expressed and methylated. We finally report the top-ranking significant gene-markerswith biological validation. Moreover, our framework improves positive predictive rate and reduces false positive rate in marker identification. In addition, we provide a comparative analysis of our gene-selection method as well as othermethods based on classificationperformances obtained using several well-known classifiers.
Witold Pedrycz - One of the best experts on this subject based on the ideXlab platform.
-
granular multi label feature selection based on mutual information
Pattern Recognition, 2017Co-Authors: Duoqian Miao, Witold PedryczAbstract:We granulate the label space into information granules to exploit label dependency.We present a multi-label maximal correlation minimal Redundancy Criterion.The proposed method can select compact and specific feature subsets.The proposed method can significantly improve the algorithm performance. Like the traditional machine learning, the multi-label learning is faced with the curse of dimensionality. Some feature selection algorithms have been proposed for multi-label learning, which either convert the multi-label feature selection problem into numerous single-label feature selection problems, or directly select features from the multi-label data set. However, the former omit the label dependency, or produce too many new labels leading to learning with significant difficulties; the latter, taking the global label dependency into consideration, usually select a few redundant or irrelevant features, because actually not all labels depend on each other, which may confuse the algorithm and degrade its classification performance. To select a more relevant and compact feature subset as well as explore the label dependency, a granular feature selection method for multi-label learning is proposed with a maximal correlation minimal Redundancy Criterion based on mutual information. The maximal correlation minimal Redundancy Criterion makes sure that the selected feature subset contains the most class-discriminative information, while in the meantime exhibits the least intra-Redundancy. Granulation can help explore the label dependency. We study the relation of the label granularity and the performance on four data sets, and compare the proposed method with other three multi-label feature selection methods. The experimental results demonstrate that the proposed method can select compact and specific feature subsets, improve the classification performance and performs better than other three methods on the widely-used multi-label learning evaluation criteria.
Youlong Yang - One of the best experts on this subject based on the ideXlab platform.
-
maximum relevance minimum common Redundancy feature selection for nonlinear data
Information Sciences, 2017Co-Authors: Jinxing Che, Youlong Yang, Xuying Bai, Shenghu Zhang, Chengzhi DengAbstract:In recent years, feature selection based on relevance Redundancy trade-off criteria has become a very promising and popular approach in the field of machine learning. However, the existing algorithmic frameworks of mutual information feature selection have certain limitations for the common feature selection problems in practice. To overcome these limitations, the idea of a new framework is developed by introducing a novel maximum relevance and minimum common Redundancy Criterion and a minimax nonlinear optimization approach. In particular, a novel mutual information feature selection method based on the normalization of the maximum relevance and minimum common Redundancy (N-MRMCR-MI) is presented, which produces a normalized value in the range [0, 1] and results in a regression problem. We perform extensive experimental comparisons over numerous state-of-art algorithms using different forecasts (Bayesian Additive Regression tree, treed Gaussian process, k-NN, and SVM) and different data sets (two simulated and five real datasets). The results show that the proposed algorithm outperforms the others in terms of feature selection and forecasting accuracy.