The Experts below are selected from a list of 48108 Experts worldwide ranked by ideXlab platform

Alex A Freitas - One of the best experts on this subject based on the ideXlab platform.

  • an artificial immune system for fuzzy Rule induction in data mining
    Lecture Notes in Computer Science, 2004
    Co-Authors: Roberto T Alves, Myriam Regattieri Delgado, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    This work proposes a Classification-Rule discovery algorithm integrating artificial immune systems and fuzzy systems. The algorithm consists of two parts: a sequential covering procedure and a Rule evolution procedure. Each antibody (candidate solution) corresponds to a Classification Rule. The Classification of new examples (antigens) considers not only the fitness of a fuzzy Rule based on the entire training set, but also the affinity between the Rule and the new example. This affinity must be greater than a threshold in order for the fuzzy Rule to be activated, and it is proposed an adaptive procedure for computing this threshold for each Rule. This paper reports results for the proposed algorithm in several data sets. Results are analyzed with respect to both predictive accuracy and Rule set simplicity, and are compared with C4.5Rules, a very popular data mining algorithm.

  • an ant colony algorithm for Classification Rule discovery
    2002
    Co-Authors: Rafael Stubs Parpinelli, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    Real life problems are known to be messy, dynamic and multi-objective, and involve high levels of uncertainty and constraints. Because traditional problem-solving methods are no longer capable of handling this level of complexity, heuristic search methods have attracted increasing attention in recent years for solving such problems. Inspired by nature, biology, statistical mechanics, physics and neuroscience, heuristics techniques are used to solve many problems where traditional methods have failed. Data Mining: A Hueristic Approach will be a repository for the applications of these techniques in the area of data mining.

  • discovering comprehensible Classification Rules with a genetic algorithm
    Congress on Evolutionary Computation, 2000
    Co-Authors: M V Fidelis, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    Presents a Classification algorithm based on genetic algorithms (GAs) that discovers comprehensible IF-THEN Rules, in the spirit of data mining. The proposed GA has a flexible chromosome encoding, where each chromosome corresponds to a Classification Rule. Although the number of genes (the genotype) is fixed, the number of Rule conditions (the phenotype) is variable. The GA also has specific mutation operators for this chromosome encoding. The algorithm was evaluated on two public-domain real-world data sets (in the medical domains of dermatology and breast cancer).

Heitor Silverio Lopes - One of the best experts on this subject based on the ideXlab platform.

  • gepclass a Classification Rule discovery tool using gene expression programming
    Lecture Notes in Computer Science, 2006
    Co-Authors: Wagner Rodrigo Weinert, Heitor Silverio Lopes
    Abstract:

    This work describes the use of a recently proposed technique - gene expression programming - for knowledge discovery in the data mining task of data Classification. We propose a new method for Rule encoding and genetic operators that preserve Rule integrity, and implemented a system, named GEPCLASS. Due to its encoding scheme, the system allows the automatic discovery of flexible Rules, better fitted to data. The performance of GEPCLASS was compared with two genetic programming systems and with C4.5, over four data sets in a five-fold cross-validation procedure. The predictive accuracy for the methods compared were similar, but the computational effort needed by GEPCLASS was significantly smaller than the other. GEPCLASS was able to find simple and accurate Rules as it can handle continuous and categorical attributes.

  • an artificial immune system for fuzzy Rule induction in data mining
    Lecture Notes in Computer Science, 2004
    Co-Authors: Roberto T Alves, Myriam Regattieri Delgado, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    This work proposes a Classification-Rule discovery algorithm integrating artificial immune systems and fuzzy systems. The algorithm consists of two parts: a sequential covering procedure and a Rule evolution procedure. Each antibody (candidate solution) corresponds to a Classification Rule. The Classification of new examples (antigens) considers not only the fitness of a fuzzy Rule based on the entire training set, but also the affinity between the Rule and the new example. This affinity must be greater than a threshold in order for the fuzzy Rule to be activated, and it is proposed an adaptive procedure for computing this threshold for each Rule. This paper reports results for the proposed algorithm in several data sets. Results are analyzed with respect to both predictive accuracy and Rule set simplicity, and are compared with C4.5Rules, a very popular data mining algorithm.

  • an ant colony algorithm for Classification Rule discovery
    2002
    Co-Authors: Rafael Stubs Parpinelli, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    Real life problems are known to be messy, dynamic and multi-objective, and involve high levels of uncertainty and constraints. Because traditional problem-solving methods are no longer capable of handling this level of complexity, heuristic search methods have attracted increasing attention in recent years for solving such problems. Inspired by nature, biology, statistical mechanics, physics and neuroscience, heuristics techniques are used to solve many problems where traditional methods have failed. Data Mining: A Hueristic Approach will be a repository for the applications of these techniques in the area of data mining.

  • discovering comprehensible Classification Rules with a genetic algorithm
    Congress on Evolutionary Computation, 2000
    Co-Authors: M V Fidelis, Heitor Silverio Lopes, Alex A Freitas
    Abstract:

    Presents a Classification algorithm based on genetic algorithms (GAs) that discovers comprehensible IF-THEN Rules, in the spirit of data mining. The proposed GA has a flexible chromosome encoding, where each chromosome corresponds to a Classification Rule. Although the number of genes (the genotype) is fixed, the number of Rule conditions (the phenotype) is variable. The GA also has specific mutation operators for this chromosome encoding. The algorithm was evaluated on two public-domain real-world data sets (in the medical domains of dermatology and breast cancer).

Enrique Vidal - One of the best experts on this subject based on the ideXlab platform.

  • learning weighted metrics to minimize nearest neighbor Classification error
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2006
    Co-Authors: Roberto Paredes, Enrique Vidal
    Abstract:

    In order to optimize the accuracy of the nearest-neighbor Classification Rule, a weighted distance is proposed, along with algorithms to automatically learn the corresponding weights. These weights may be specific for each class and feature, for each individual prototype, or for both. The learning algorithms are derived by (approximately) minimizing the leaving-one-out Classification error of the given training set. The proposed approach is assessed through a series of experiments with UCI/STATLOG corpora, as well as with a more specific task of text Classification which entails very sparse data representation and huge dimensionality. In all these experiments, the proposed approach shows a uniformly good behavior, with results comparable to or better than state-of-the-art results published with the same data so far

Linjun Zhang - One of the best experts on this subject based on the ideXlab platform.

  • high dimensional linear discriminant analysis optimality adaptive algorithm and missing data
    Journal of The Royal Statistical Society Series B-statistical Methodology, 2019
    Co-Authors: Tommaso Cai, Linjun Zhang
    Abstract:

    The paper develops optimality theory for linear discriminant analysis in the high dimensional setting. A data‐driven and tuning‐free Classification Rule, which is based on an adaptive constrained l1‐minimization approach, is proposed and analysed. Minimax lower bounds are obtained and this Classification Rule is shown to be simultaneously rate optimal over a collection of parameter spaces. In addition, we consider Classification with incomplete data under the missingness completely at random model. An adaptive classifier with theoretical guarantees is introduced and the optimal rate of convergence for high dimensional linear discriminant analysis under the missingness completely at random model is established. The technical analysis for the case of missing data is much more challenging than that for complete data. We establish a large deviation result for the generalized sample covariance matrix, which serves as a key technical tool and can be of independent interest. An application to lung cancer and leukaemia studies is also discussed.

  • high dimensional linear discriminant analysis optimality adaptive algorithm and missing data
    arXiv: Methodology, 2018
    Co-Authors: Tommaso Cai, Linjun Zhang
    Abstract:

    This paper aims to develop an optimality theory for linear discriminant analysis in the high-dimensional setting. A data-driven and tuning free Classification Rule, which is based on an adaptive constrained $\ell_1$ minimization approach, is proposed and analyzed. Minimax lower bounds are obtained and this Classification Rule is shown to be simultaneously rate optimal over a collection of parameter spaces. In addition, we consider Classification with incomplete data under the missing completely at random (MCR) model. An adaptive classifier with theoretical guarantees is introduced and optimal rate of convergence for high-dimensional linear discriminant analysis under the MCR model is established. The technical analysis for the case of missing data is much more challenging than that for the complete data. We establish a large deviation result for the generalized sample covariance matrix, which serves as a key technical tool and can be of independent interest. An application to lung cancer and leukemia studies is also discussed.

Niall Rooney - One of the best experts on this subject based on the ideXlab platform.

  • temporal data mining for smart homes
    Lecture Notes in Computer Science, 2006
    Co-Authors: Mykola Galushka, Dave Patterson, Niall Rooney
    Abstract:

    Temporal data mining is a relatively new area of research in computer science. It can provide a large variety of different methods and techniques for handling and analyzing temporal data generated by smart-home environments. Temporal data mining in general fits into a two level architecture, where initially a transformation technique reduces data dimensionality in the first level and indexing techniques provide efficient access to the data in the second level. This infrastructure of temporal data mining provides the basis for high-level data mining operations such as clustering, Classification, Rule discovery and prediction. These operations can form the basis for developing different smart-home applications, capable of addressing a number of situations occurring within this environment. This paper outlines the main temporal data mining techniques available and provides examples of where they can be applied within a smart home environment.