The Experts below are selected from a list of 114054 Experts worldwide ranked by ideXlab platform
Christopher Yau - One of the best experts on this subject based on the ideXlab platform.
-
zifa Dimensionality Reduction for zero inflated single cell gene expression analysis
Genome Biology, 2015Co-Authors: Emma Pierson, Christopher YauAbstract:Single-cell RNA-seq data allows insight into normal cellular function and various disease states through molecular characterization of gene expression on the single cell level. Dimensionality Reduction of such high-dimensional data sets is essential for visualization and analysis, but single-cell RNA-seq data are challenging for classical Dimensionality-Reduction methods because of the prevalence of dropout events, which lead to zero-inflated data. Here, we develop a Dimensionality-Reduction method, (Z)ero (I)nflated (F)actor (A)nalysis (ZIFA), which explicitly models the dropout characteristics, and show that it improves modeling accuracy on simulated and biological data sets.
-
Dimensionality Reduction for zero inflated single cell gene expression analysis
bioRxiv, 2015Co-Authors: Christopher Yau, Emma PiersonAbstract:Single cell RNA-seq data allows insight into normal cellular function and diseases including cancer through the molecular characterisation of cellular state at the single-cell level. Dimensionality Reduction of such high-dimensional datasets is essential for visualization and analysis, but single-cell RNA-seq data is challenging for classical Dimensionality Reduction methods because of the prevalence of dropout events leading to zero-inflated data. Here we develop a Dimensionality Reduction method, (Z)ero (I)nflated (F)actor (A)nalysis (ZIFA), which explicitly models the dropout characteristics, and show that it improves performance on simulated and biological datasets.
Michel Verleysen - One of the best experts on this subject based on the ideXlab platform.
-
Nonlinear Dimensionality Reduction for visualization
Lecture Notes in Computer Science, 2013Co-Authors: Michel Verleysen, John Aldo LeeAbstract:The visual interpretation of data is an essential step to guide any further processing or decision making. Dimensionality Reduction (or manifold learning) tools may be used for visualization if the resulting dimension is constrained to be 2 or 3. The field of machine learning has developed numerous nonlinear Dimensionality Reduction tools in the last decades. However, the diversity of methods reflects the diversity of quality criteria used both for optimizing the algorithms, and for assessing their performances. In addition, these criteria are not always compatible with subjective visual quality. Finally, the Dimensionality Reduction methods themselves do not always possess computational properties that are compatible with interactive data visualization. This paper presents current and future developments to use Dimensionality Reduction methods for data visualization.
-
IJCNN - Unsupervised Dimensionality Reduction: Overview and recent advances
The 2010 International Joint Conference on Neural Networks (IJCNN), 2010Co-Authors: John Aldo Lee, Michel VerleysenAbstract:Unsupervised Dimensionality Reduction aims at representing high-dimensional data in lower-dimensional spaces in a faithful way. Dimensionality Reduction can be used for compression or denoising purposes, but data visualization remains one its most prominent applications. This paper attempts to give a broad overview of the domain. Past develoments are briefly introduced and pinned up on the time line of the last eleven decades. Next, the principles and techniques involved in the major methods are described. A taxonomy of the methods is suggested, taking into account various properties. Finally, the issue of quality assessment is briefly dealt with.
-
nonlinear Dimensionality Reduction
2007Co-Authors: John Aldo Lee, Michel VerleysenAbstract:Methods of Dimensionality Reduction provide a way to understand and visualize the structure of complex data sets. Traditional methods like principal component analysis and classical metric multidimensional scaling suffer from being based on linear models. Until recently, very few methods were able to reduce the data Dimensionality in a nonlinear way. However, since the late nineties, many new methods have been developed and nonlinear Dimensionality Reduction, also called manifold learning, has become a hot topic. New advances that account for this rapid growth are, e.g. the use of graphs to represent the manifold topology, and the use of new metrics like the geodesic distance. In addition, new optimization schemes, based on kernel techniques and spectral decomposition, have lead to spectral embedding, which encompasses many of the recently developed methods. This book describes existing and advanced methods to reduce the Dimensionality of numerical databases. For each method, the description starts from intuitive ideas, develops the necessary mathematical details, and ends by outlining the algorithmic implementation. Methods are compared with each other with the help of different illustrative examples. The purpose of the book is to summarize clear facts and ideas about well-known methods as well as recent developments in the topic of nonlinear Dimensionality Reduction. With this goal in mind, methods are all described from a unifying point of view, in order to highlight their respective strengths and shortcomings. The book is primarily intended for statisticians, computer scientists and data analysts. It is also accessible to other practitioners having a basic background in statistics and/or computational learning, like psychologists (in psychometry) and economists.
John Aldo Lee - One of the best experts on this subject based on the ideXlab platform.
-
Nonlinear Dimensionality Reduction for visualization
Lecture Notes in Computer Science, 2013Co-Authors: Michel Verleysen, John Aldo LeeAbstract:The visual interpretation of data is an essential step to guide any further processing or decision making. Dimensionality Reduction (or manifold learning) tools may be used for visualization if the resulting dimension is constrained to be 2 or 3. The field of machine learning has developed numerous nonlinear Dimensionality Reduction tools in the last decades. However, the diversity of methods reflects the diversity of quality criteria used both for optimizing the algorithms, and for assessing their performances. In addition, these criteria are not always compatible with subjective visual quality. Finally, the Dimensionality Reduction methods themselves do not always possess computational properties that are compatible with interactive data visualization. This paper presents current and future developments to use Dimensionality Reduction methods for data visualization.
-
IJCNN - Unsupervised Dimensionality Reduction: Overview and recent advances
The 2010 International Joint Conference on Neural Networks (IJCNN), 2010Co-Authors: John Aldo Lee, Michel VerleysenAbstract:Unsupervised Dimensionality Reduction aims at representing high-dimensional data in lower-dimensional spaces in a faithful way. Dimensionality Reduction can be used for compression or denoising purposes, but data visualization remains one its most prominent applications. This paper attempts to give a broad overview of the domain. Past develoments are briefly introduced and pinned up on the time line of the last eleven decades. Next, the principles and techniques involved in the major methods are described. A taxonomy of the methods is suggested, taking into account various properties. Finally, the issue of quality assessment is briefly dealt with.
-
nonlinear Dimensionality Reduction
2007Co-Authors: John Aldo Lee, Michel VerleysenAbstract:Methods of Dimensionality Reduction provide a way to understand and visualize the structure of complex data sets. Traditional methods like principal component analysis and classical metric multidimensional scaling suffer from being based on linear models. Until recently, very few methods were able to reduce the data Dimensionality in a nonlinear way. However, since the late nineties, many new methods have been developed and nonlinear Dimensionality Reduction, also called manifold learning, has become a hot topic. New advances that account for this rapid growth are, e.g. the use of graphs to represent the manifold topology, and the use of new metrics like the geodesic distance. In addition, new optimization schemes, based on kernel techniques and spectral decomposition, have lead to spectral embedding, which encompasses many of the recently developed methods. This book describes existing and advanced methods to reduce the Dimensionality of numerical databases. For each method, the description starts from intuitive ideas, develops the necessary mathematical details, and ends by outlining the algorithmic implementation. Methods are compared with each other with the help of different illustrative examples. The purpose of the book is to summarize clear facts and ideas about well-known methods as well as recent developments in the topic of nonlinear Dimensionality Reduction. With this goal in mind, methods are all described from a unifying point of view, in order to highlight their respective strengths and shortcomings. The book is primarily intended for statisticians, computer scientists and data analysts. It is also accessible to other practitioners having a basic background in statistics and/or computational learning, like psychologists (in psychometry) and economists.
Zhihua Zhou - One of the best experts on this subject based on the ideXlab platform.
-
multilabel Dimensionality Reduction via dependence maximization
ACM Transactions on Knowledge Discovery From Data, 2010Co-Authors: Yin Zhang, Zhihua ZhouAbstract:Multilabel learning deals with data associated with multiple labels simultaneously. Like other data mining and machine learning tasks, multilabel learning also suffers from the curse of Dimensionality. Dimensionality Reduction has been studied for many years, however, multilabel Dimensionality Reduction remains almost untouched. In this article, we propose a multilabel Dimensionality Reduction method, MDDM, with two kinds of projection strategies, attempting to project the original data into a lower-dimensional feature space maximizing the dependence between the original feature description and the associated class labels. Based on the Hilbert-Schmidt Independence Criterion, we derive a eigen-decomposition problem which enables the Dimensionality Reduction process to be efficient. Experiments validate the performance of MDDM.
-
AAAI - Multi-instance Dimensionality Reduction
2010Co-Authors: Yuyin Sun, Zhihua ZhouAbstract:Multi-instance learning deals with problems that treat bags of instances as training examples. In single-instance learning problems, Dimensionality Reduction is an essential step for high-dimensional data analysis and has been studied for years. The curse of Dimensionality also exists in multi-instance learning tasks, yet this difficult task has not been studied before. Direct application of existing single-instance Dimensionality Reduction objectives to multi-instance learning tasks may not work well since it ignores the characteristic of multi-instance learning that the labels of bags are known while the labels of instances are unknown. In this paper, we propose an effective model and develop an efficient algorithm to solve the multi-instance Dimensionality Reduction problem. We formulate the objective as an optimization problem by considering orthonormality and sparsity constraints in the projection matrix for Dimensionality Reduction, and then solve it by the gradient descent along the tangent space of the orthonormal matrices. We also propose an approximation for improving the efficiency. Experimental results validate the effectiveness of the proposed method.
-
semi supervised Dimensionality Reduction
SIAM International Conference on Data Mining, 2007Co-Authors: Daoqiang Zhang, Zhihua Zhou, Songcan ChenAbstract:Dimensionality Reduction is among the keys in mining highdimensional data. This paper studies semi-supervised Dimensionality Reduction. In this setting, besides abundant unlabeled examples, domain knowledge in the form of pairwise constraints are available, which specifies whether a pair of instances belong to the same class (must-link constraints) or different classes (cannot-link constraints). We propose the SSDR algorithm, which can preserve the intrinsic structure of the unlabeled data as well as both the must-link and cannot-link constraints defined on the labeled examples in the projected low-dimensional space. The SSDR algorithm is efficient and has a closed form solution. Experiments on a broad range of data sets show that SSDR is superior to many established Dimensionality Reduction methods.
-
supervised nonlinear Dimensionality Reduction for visualization and classification
Systems Man and Cybernetics, 2005Co-Authors: Xin Geng, Dechuan Zhan, Zhihua ZhouAbstract:When performing visualization and classification, people often confront the problem of Dimensionality Reduction. Isomap is one of the most promising nonlinear Dimensionality Reduction techniques. However, when Isomap is applied to real-world data, it shows some limitations, such as being sensitive to noise. In this paper, an improved version of Isomap, namely S-Isomap, is proposed. S-Isomap utilizes class information to guide the procedure of nonlinear Dimensionality Reduction. Such a kind of procedure is called supervised nonlinear Dimensionality Reduction. In S-Isomap, the neighborhood graph of the input data is constructed according to a certain kind of dissimilarity between data points, which is specially designed to integrate the class information. The dissimilarity has several good properties which help to discover the true neighborhood of the data and, thus, makes S-Isomap a robust technique for both visualization and classification, especially for real-world problems. In the visualization experiments, S-Isomap is compared with Isomap, LLE, and WeightedIso. The results show that S-Isomap performs the best. In the classification experiments, S-Isomap is used as a preprocess of classification and compared with Isomap, WeightedIso, as well as some other well-established classification methods, including the K-nearest neighbor classifier, BP neural network, J4.8 decision tree, and SVM. The results reveal that S-Isomap excels compared to Isomap and WeightedIso in classification, and it is highly competitive with those well-known classification methods.
Masashi Sugiyama - One of the best experts on this subject based on the ideXlab platform.
-
Nonlinear Dimensionality Reduction
Introduction to Statistical Machine Learning, 2016Co-Authors: Masashi SugiyamaAbstract:In this chapter, supervised and unsupervised methods of nonlinear Dimensionality Reduction are introduced, including approaches based on kernels and neural networks.
-
Dimensionality Reduction of multimodal labeled data by local fisher discriminant analysis
Journal of Machine Learning Research, 2007Co-Authors: Masashi SugiyamaAbstract:Reducing the Dimensionality of data without losing intrinsic information is an important preprocessing step in high-dimensional data analysis. Fisher discriminant analysis (FDA) is a traditional technique for supervised Dimensionality Reduction, but it tends to give undesired results if samples in a class are multimodal. An unsupervised Dimensionality Reduction method called locality-preserving projection (LPP) can work well with multimodal data due to its locality preserving property. However, since LPP does not take the label information into account, it is not necessarily useful in supervised learning scenarios. In this paper, we propose a new linear supervised Dimensionality Reduction method called local Fisher discriminant analysis (LFDA), which effectively combines the ideas of FDA and LPP. LFDA has an analytic form of the embedding transformation and the solution can be easily computed just by solving a generalized eigenvalue problem. We demonstrate the practical usefulness and high scalability of the LFDA method in data visualization and classification tasks through extensive simulation studies. We also show that LFDA can be extended to non-linear Dimensionality Reduction scenarios by applying the kernel trick.
-
local fisher discriminant analysis for supervised Dimensionality Reduction
International Conference on Machine Learning, 2006Co-Authors: Masashi SugiyamaAbstract:Dimensionality Reduction is one of the important preprocessing steps in high-dimensional data analysis. In this paper, we consider the supervised Dimensionality Reduction problem where samples are accompanied with class labels. Traditional Fisher discriminant analysis is a popular and powerful method for this purpose. However, it tends to give undesired results if samples in some class form several separate clusters, i.e., multimodal. In this paper, we propose a new Dimensionality Reduction method called local Fisher discriminant analysis (LFDA), which is a localized variant of Fisher discriminant analysis. LFDA takes local structure of the data into account so the multimodal data can be embedded appropriately. We also show that LFDA can be extended to non-linear Dimensionality Reduction scenarios by the kernel trick.