The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform
Trevor Hastie - One of the best experts on this subject based on the ideXlab platform.
-
Classification of gene microarrays by penalized logistic regression
Biostatistics, 2004Co-Authors: Ji Zhu, Trevor HastieAbstract:Classification of patient samples is an important aspect of cancer diagnosis and treatment. The support vector machine (SVM) has been successfully applied to microarray cancer diagnosis problems. However, one weakness of the SVM is that given a tumor sample, it only predicts a cancer Class label but does not provide any estimate of the underlying probability. We propose penalized logistic regression (PLR) as an alternative to the SVM for the microarray cancer diagnosis problem. We show that when using the same set of genes, PLR and the SVM perform similarly in cancer Classification, but PLR has the advantage of additionally providing an estimate of the underlying probability. Often a primary goal in microarray cancer diagnosis is to identify the genes responsible for the Classification, rather than Class Prediction. We consider two gene selection methods in this paper, univariate ranking (UR) and recursive feature elimination (RFE). Empirical results indicate that PLR combined with RFE tends to select fewer genes than other methods and also performs well in both cross-validation and test samples. A fast algorithm for solving PLR is also described.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to the other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
diagnosis of multiple cancer types by shrunken centroids of gene expression
Proceedings of the National Academy of Sciences of the United States of America, 2002Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian NarasimhanAbstract:We have devised an approach to cancer Class Prediction from gene expression profiling, based on an enhancement of the simple nearest prototype (centroid) Classifier. We shrink the prototypes and hence obtain a Classifier that is often more accurate than competing methods. Our method of “nearest shrunken centroids” identifies subsets of genes that best characterize each Class. The technique is general and can be used in many other Classification problems. To demonstrate its effectiveness, we show that the method was highly efficient in finding genes for Classifying small round blue cell tumors and leukemias.
Xiaoqi Zheng - One of the best experts on this subject based on the ideXlab platform.
-
Zheng X: PSSP-RFE: accurate Prediction of protein structural Class by recursive feature extraction from PSI-BLAST profile, physical-chemical property and functional annotations. PLoS One 2014
2016Co-Authors: Xiang Cui, Hua Yang, Yue Zhou, Yuan Zhang, Zhong Luo, Xiaoqi ZhengAbstract:Protein structure Prediction is critical to functional annotation of the massively accumulated biological sequences, which prompts an imperative need for the development of high-throughput technologies. As a first and key step in protein structure Prediction, protein structural Class Prediction becomes an increasingly challenging task. Amongst most homological-based approaches, the accuracies of protein structural Class Prediction are sufficiently high for high similarity datasets, but still far from being satisfactory for low similarity datasets, i.e., below 40 % in pairwise sequence similarity. Therefore, we present a novel method for accurate and reliable protein structural Class Prediction for both high and low similarity datasets. This method is based on Support Vector Machine (SVM) in conjunction with integrated features from position-specific score matrix (PSSM), PROFEAT and Gene Ontology (GO). A feature selection approach, SVM-RFE, is also used to rank the integrated feature vectors through recursively removing the feature with the lowest ranking score. The definitive top features selected by SVM-RFE are input into the SVM engines to predict the structural Class of a query protein. To validate our method, jackknife tests were applied to seven widely used benchmark datasets, reaching overall accuracie
-
pssp rfe accurate Prediction of protein structural Class by recursive feature extraction from psi blast profile physical chemical property and functional annotations
PLOS ONE, 2014Co-Authors: Liqi Li, Sanjiu Yu, Xiaoqi Zheng, Hua Yang, Yue Zhou, Yuan ZhangAbstract:Protein structure Prediction is critical to functional annotation of the massively accumulated biological sequences, which prompts an imperative need for the development of high-throughput technologies. As a first and key step in protein structure Prediction, protein structural Class Prediction becomes an increasingly challenging task. Amongst most homological-based approaches, the accuracies of protein structural Class Prediction are sufficiently high for high similarity datasets, but still far from being satisfactory for low similarity datasets, i.e., below 40% in pairwise sequence similarity. Therefore, we present a novel method for accurate and reliable protein structural Class Prediction for both high and low similarity datasets. This method is based on Support Vector Machine (SVM) in conjunction with integrated features from position-specific score matrix (PSSM), PROFEAT and Gene Ontology (GO). A feature selection approach, SVM-RFE, is also used to rank the integrated feature vectors through recursively removing the feature with the lowest ranking score. The definitive top features selected by SVM-RFE are input into the SVM engines to predict the structural Class of a query protein. To validate our method, jackknife tests were applied to seven widely used benchmark datasets, reaching overall accuracies between 84.61% and 99.79%, which are significantly higher than those achieved by state-of-the-art tools. These results suggest that our method could serve as an accurate and cost-effective alternative to existing methods in protein structural Classification, especially for low similarity datasets.
-
accurate Prediction of protein structural Class using auto covariance transformation of psi blast profiles
Amino Acids, 2012Co-Authors: Taigang Liu, Xiaoqi Zheng, Xingbo Geng, Jun WangAbstract:Computational Prediction of protein structural Class based solely on sequence data remains a challenging problem in protein science. Existing methods differ in the protein sequence representation models and Prediction engines adopted. In this study, a powerful feature extraction method, which combines position-specific score matrix (PSSM) with auto covariance (AC) transformation, is introduced. Thus, a sample protein is represented by a series of discrete components, which could partially incorporate the long-range sequence order information and evolutionary information reflected from the PSI-BLAST profile. To verify the performance of our method, jackknife cross-validation tests are performed on four widely used benchmark datasets. Comparison of our results with existing methods shows that our method provides the state-of-the-art performance for structural Class Prediction. A Web server that implements the proposed method is freely available at http://202.194.133.5/xinxi/AAC_PSSM_AC/index.htm.
Balasubramanian Narasimhan - One of the best experts on this subject based on the ideXlab platform.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to the other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
diagnosis of multiple cancer types by shrunken centroids of gene expression
Proceedings of the National Academy of Sciences of the United States of America, 2002Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian NarasimhanAbstract:We have devised an approach to cancer Class Prediction from gene expression profiling, based on an enhancement of the simple nearest prototype (centroid) Classifier. We shrink the prototypes and hence obtain a Classifier that is often more accurate than competing methods. Our method of “nearest shrunken centroids” identifies subsets of genes that best characterize each Class. The technique is general and can be used in many other Classification problems. To demonstrate its effectiveness, we show that the method was highly efficient in finding genes for Classifying small round blue cell tumors and leukemias.
Robert Tibshirani - One of the best experts on this subject based on the ideXlab platform.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to the other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
diagnosis of multiple cancer types by shrunken centroids of gene expression
Proceedings of the National Academy of Sciences of the United States of America, 2002Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian NarasimhanAbstract:We have devised an approach to cancer Class Prediction from gene expression profiling, based on an enhancement of the simple nearest prototype (centroid) Classifier. We shrink the prototypes and hence obtain a Classifier that is often more accurate than competing methods. Our method of “nearest shrunken centroids” identifies subsets of genes that best characterize each Class. The technique is general and can be used in many other Classification problems. To demonstrate its effectiveness, we show that the method was highly efficient in finding genes for Classifying small round blue cell tumors and leukemias.
Gilbert Chu - One of the best experts on this subject based on the ideXlab platform.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.
-
Class Prediction by nearest shrunken centroids with applications to dna microarrays
Statistical Science, 2003Co-Authors: Robert Tibshirani, Trevor Hastie, Balasubramanian Narasimhan, Gilbert ChuAbstract:We propose a new method for Class Prediction in DNA microarray studies based on an enhancement of the nearest prototype Classifier. Our technique uses "shrunken" centroids as prototypes for each Class to identify the subsets of the genes that best characterize each Class. The method is general and can be applied to the other high-dimensional Classification problems. The method is illustrated on data from two gene expression studies: lymphoma and cancer cell lines.