The Experts below are selected from a list of 18159 Experts worldwide ranked by ideXlab platform

Alex A. Freitas - One of the best experts on this subject based on the ideXlab platform.

  • PKDD - A Genetic Algorithm-Based Solution for the Problem of Small Disjuncts
    Principles of Data Mining and Knowledge Discovery, 2020
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    In essence, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. The problem is particularly serious because, although each small Disjuncts covers few examples, the set of small Disjuncts can cover a large number of examples. This paper proposes a solution to the problem of discovering accurate small-disjunct rules based on genetic algorithms. The basic idea of our method is to use a hybrid decision tree / genetic algorithm approach for classification. More precisely, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm, while examples belonging to small Disjuncts are classified by a new genetic algorithm, particularly designed for discovering small-disjunct rules.

  • A hybrid decision tree/genetic algorithm method for data mining
    Information Sciences, 2004
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the objective is to predict the class which an example belongs to. Discovered knowledge is expressed in the form of high-level, easy-to-interpret classification rules. In order to discover classification rules, we propose a hybrid decision tree/genetic algorithm method. The central idea of this hybrid method involves the concept of small Disjuncts in data mining, as follows. In essence, a set of classification rules can be regarded as a logical disjunction of rules, so that each rule can be regarded as a disjunct. A small disjunct is a rule covering a small number of examples. Due to their nature, small Disjuncts are error prone. However, although each small disjunct covers just a few examples, the set of all small Disjuncts can cover a large number of examples, so that it is important to develop new approaches to cope with the problem of small Disjuncts. In our hybrid approach, we have developed two genetic algorithms (GA) specifically designed for discovering rules covering examples belonging to small Disjuncts, whereas a conventional decision tree algorithm is used to produce rules covering examples belonging to large Disjuncts. We present results evaluating the performance of the hybrid method in 22 real-world data sets. © 2003 Elsevier Inc. All rights reserved.

  • A genetic-algorithm for discovering small-disjunct rules in data mining
    Applied Soft Computing, 2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the goal is to discover rules predicting the class of examples (records of a given dataset). In the context of data mining, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. At first glance, this is not a serious problem, since the impact on predictive accuracy should be small. However, although each small-disjunct covers few examples, the set of all small Disjuncts can cover a large number of examples. This paper presents evidence that this is the case in several datasets. This paper also addresses the problem of small Disjuncts by using a hybrid decision-tree/genetic-algorithm approach. In essence, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic-algorithm specifically designed for discovering small-disjunct rules. We present results comparing the predictive accuracy of this hybrid system with the prediction accuracy of three versions of C4.5 alone in eight public domain datasets. Overall, the results show that our hybrid system achieves better predictive accuracy than all three versions of C4.5 alone.

  • GECCO - A genetic algorithm with sequential niching for discovering small-disjunct rules
    2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This work addresses the well-known classification task of data mining. In this context, small Disjuncts are classification rules covering a small number of examples. One approach for coping with small Disjuncts, proposed in our previous work, consists of using a decision-tree/genetic algorithm method. The basic idea is that examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic algorithm (GA) designed for discovering small-disjunct rules. In this paper we follow this basic idea, but we propose a new GA which consists of several major modifications to the original GA used for coping with small Disjuncts. The performance of the new GA is extensively evaluated by comparing it with two versions of C4.5, across several data sets, and with several different sizes of small Disjuncts.

  • A genetic-algorithm for discovering small-disjunct rules in data mining
    Applied Soft Computing Journal, 2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the goal is to discover rules predicting the class of examples (records of a given dataset). In the context of data mining, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. At first glance, this is not a serious problem, since the impact on predictive accuracy should be small. However, although each small-disjunct covers few examples, the set of all small Disjuncts can cover a large number of examples. This paper presents evidence that this is the case in several datasets. This paper also addresses the problem of small Disjuncts by using a hybrid decision-tree/genetic-algorithm approach. In essence, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic-algorithm specifically designed for discovering small-disjunct rules. We present results comparing the predictive accuracy of this hybrid system with the prediction accuracy of three versions of C4.5 alone in eight public domain datasets. Overall, the results show that our hybrid system achieves better predictive accuracy than all three versions of C4.5 alone. ?? 2002 Elsevier Science B.V. All rights reserved.

Deborah R. Carvalho - One of the best experts on this subject based on the ideXlab platform.

  • PKDD - A Genetic Algorithm-Based Solution for the Problem of Small Disjuncts
    Principles of Data Mining and Knowledge Discovery, 2020
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    In essence, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. The problem is particularly serious because, although each small Disjuncts covers few examples, the set of small Disjuncts can cover a large number of examples. This paper proposes a solution to the problem of discovering accurate small-disjunct rules based on genetic algorithms. The basic idea of our method is to use a hybrid decision tree / genetic algorithm approach for classification. More precisely, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm, while examples belonging to small Disjuncts are classified by a new genetic algorithm, particularly designed for discovering small-disjunct rules.

  • A hybrid decision tree/genetic algorithm method for data mining
    Information Sciences, 2004
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the objective is to predict the class which an example belongs to. Discovered knowledge is expressed in the form of high-level, easy-to-interpret classification rules. In order to discover classification rules, we propose a hybrid decision tree/genetic algorithm method. The central idea of this hybrid method involves the concept of small Disjuncts in data mining, as follows. In essence, a set of classification rules can be regarded as a logical disjunction of rules, so that each rule can be regarded as a disjunct. A small disjunct is a rule covering a small number of examples. Due to their nature, small Disjuncts are error prone. However, although each small disjunct covers just a few examples, the set of all small Disjuncts can cover a large number of examples, so that it is important to develop new approaches to cope with the problem of small Disjuncts. In our hybrid approach, we have developed two genetic algorithms (GA) specifically designed for discovering rules covering examples belonging to small Disjuncts, whereas a conventional decision tree algorithm is used to produce rules covering examples belonging to large Disjuncts. We present results evaluating the performance of the hybrid method in 22 real-world data sets. © 2003 Elsevier Inc. All rights reserved.

  • A genetic-algorithm for discovering small-disjunct rules in data mining
    Applied Soft Computing, 2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the goal is to discover rules predicting the class of examples (records of a given dataset). In the context of data mining, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. At first glance, this is not a serious problem, since the impact on predictive accuracy should be small. However, although each small-disjunct covers few examples, the set of all small Disjuncts can cover a large number of examples. This paper presents evidence that this is the case in several datasets. This paper also addresses the problem of small Disjuncts by using a hybrid decision-tree/genetic-algorithm approach. In essence, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic-algorithm specifically designed for discovering small-disjunct rules. We present results comparing the predictive accuracy of this hybrid system with the prediction accuracy of three versions of C4.5 alone in eight public domain datasets. Overall, the results show that our hybrid system achieves better predictive accuracy than all three versions of C4.5 alone.

  • GECCO - A genetic algorithm with sequential niching for discovering small-disjunct rules
    2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This work addresses the well-known classification task of data mining. In this context, small Disjuncts are classification rules covering a small number of examples. One approach for coping with small Disjuncts, proposed in our previous work, consists of using a decision-tree/genetic algorithm method. The basic idea is that examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic algorithm (GA) designed for discovering small-disjunct rules. In this paper we follow this basic idea, but we propose a new GA which consists of several major modifications to the original GA used for coping with small Disjuncts. The performance of the new GA is extensively evaluated by comparing it with two versions of C4.5, across several data sets, and with several different sizes of small Disjuncts.

  • A genetic-algorithm for discovering small-disjunct rules in data mining
    Applied Soft Computing Journal, 2002
    Co-Authors: Deborah R. Carvalho, Alex A. Freitas
    Abstract:

    This paper addresses the well-known classification task of data mining, where the goal is to discover rules predicting the class of examples (records of a given dataset). In the context of data mining, small Disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. At first glance, this is not a serious problem, since the impact on predictive accuracy should be small. However, although each small-disjunct covers few examples, the set of all small Disjuncts can cover a large number of examples. This paper presents evidence that this is the case in several datasets. This paper also addresses the problem of small Disjuncts by using a hybrid decision-tree/genetic-algorithm approach. In essence, examples belonging to large Disjuncts are classified by rules produced by a decision-tree algorithm (C4.5), while examples belonging to small Disjuncts are classified by a genetic-algorithm specifically designed for discovering small-disjunct rules. We present results comparing the predictive accuracy of this hybrid system with the prediction accuracy of three versions of C4.5 alone in eight public domain datasets. Overall, the results show that our hybrid system achieves better predictive accuracy than all three versions of C4.5 alone. ?? 2002 Elsevier Science B.V. All rights reserved.

Gary M Weiss - One of the best experts on this subject based on the ideXlab platform.

  • the impact of small Disjuncts on classifier learning
    Data Mining, 2010
    Co-Authors: Gary M Weiss
    Abstract:

    Many classifier induction systems express the induced classifier in terms of a disjunctive description. Small Disjuncts are those that classify few training examples. These Disjuncts are interesting because they are known to have a much higher error rate than large Disjuncts and are responsible for many, if not most, of all classification errors. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of data sets. In this chapter we provide a much more systematic study of small Disjuncts and analyze how they affect classifiers induced from 30 real-world data sets. A new metric, error concentration, is used to show that for these 30 data sets classification errors are often heavily concentrated toward the smaller Disjuncts. Various factors, including pruning, training set size, noise, and class imbalance are then analyzed to determine how they affect small Disjuncts and the distribution of errors across Disjuncts. This analysis provides many insights into why some data sets are difficult to learn from and also provides a better understanding of classifier learning in general.We believe that such an understanding is critical to the development of improved classifier induction algorithms.

  • Data Mining - The Impact of Small Disjuncts on Classifier Learning
    Annals of Information Systems, 2009
    Co-Authors: Gary M Weiss
    Abstract:

    Many classifier induction systems express the induced classifier in terms of a disjunctive description. Small Disjuncts are those that classify few training examples. These Disjuncts are interesting because they are known to have a much higher error rate than large Disjuncts and are responsible for many, if not most, of all classification errors. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of data sets. In this chapter we provide a much more systematic study of small Disjuncts and analyze how they affect classifiers induced from 30 real-world data sets. A new metric, error concentration, is used to show that for these 30 data sets classification errors are often heavily concentrated toward the smaller Disjuncts. Various factors, including pruning, training set size, noise, and class imbalance are then analyzed to determine how they affect small Disjuncts and the distribution of errors across Disjuncts. This analysis provides many insights into why some data sets are difficult to learn from and also provides a better understanding of classifier learning in general.We believe that such an understanding is critical to the development of improved classifier induction algorithms.

  • the effect of small Disjuncts and class distribution on decision tree learning
    2003
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    The main goal of classifier learning is to generate a model that makes few misclassification errors. Given this emphasis on error minimization, it makes sense to try to understand how the induction process gives rise to classifiers that make errors and whether we can identify those parts of the classifier that generate most of the errors. In this thesis we provide the first comprehensive studies of two major sources of classification errors. The first study concerns small Disjuncts, which are those Disjuncts within a classifier that cover only a few training examples. An analysis of classifiers induced from thirty data sets shows that these small Disjuncts are extremely error prone and often account for the majority of all classification errors. Because small Disjuncts largely determine classifier performance, we use them as a “lens” through which to study classifier induction. Factors such as pruning, training-set size, noise and class imbalance are each analyzed to determine how they affect small Disjuncts and, more generally, classifier learning. The second study analyzes the effect that rare classes and class distribution have on learning. Those examples belonging to rare classes are shown to be misclassified much more often than common classes. The thesis then goes on to analyze the impact that varying the class distribution of the training data has on classifier performance. The experimental results indicate that the naturally occurring class distribution is not always best for learning and that a balanced class distribution should be chosen to generate a classifier robust to different misclassification costs. It is often necessary to limit the amount of training data used for learning, due to the costs associated with obtaining and learning from this data. This thesis presents a budget-sensitive progressive-sampling algorithm for selecting training examples in this situation. This algorithm is shown to produce a class distribution that performs quite well for learning (i.e., is near optimal).

  • a quantitative study of small Disjuncts
    National Conference on Artificial Intelligence, 2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

  • AAAI/IAAI - A Quantitative Study of Small Disjuncts
    2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

Haym Hirsh - One of the best experts on this subject based on the ideXlab platform.

  • the effect of small Disjuncts and class distribution on decision tree learning
    2003
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    The main goal of classifier learning is to generate a model that makes few misclassification errors. Given this emphasis on error minimization, it makes sense to try to understand how the induction process gives rise to classifiers that make errors and whether we can identify those parts of the classifier that generate most of the errors. In this thesis we provide the first comprehensive studies of two major sources of classification errors. The first study concerns small Disjuncts, which are those Disjuncts within a classifier that cover only a few training examples. An analysis of classifiers induced from thirty data sets shows that these small Disjuncts are extremely error prone and often account for the majority of all classification errors. Because small Disjuncts largely determine classifier performance, we use them as a “lens” through which to study classifier induction. Factors such as pruning, training-set size, noise and class imbalance are each analyzed to determine how they affect small Disjuncts and, more generally, classifier learning. The second study analyzes the effect that rare classes and class distribution have on learning. Those examples belonging to rare classes are shown to be misclassified much more often than common classes. The thesis then goes on to analyze the impact that varying the class distribution of the training data has on classifier performance. The experimental results indicate that the naturally occurring class distribution is not always best for learning and that a balanced class distribution should be chosen to generate a classifier robust to different misclassification costs. It is often necessary to limit the amount of training data used for learning, due to the costs associated with obtaining and learning from this data. This thesis presents a budget-sensitive progressive-sampling algorithm for selecting training examples in this situation. This algorithm is shown to produce a class distribution that performs quite well for learning (i.e., is near optimal).

  • a quantitative study of small Disjuncts
    National Conference on Artificial Intelligence, 2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

  • AAAI/IAAI - A Quantitative Study of Small Disjuncts
    2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

  • a quantitative study of small Disjuncts experiments and results
    2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

  • A Quantitative Study of Small Disjuncts: Experiments and Results *
    2000
    Co-Authors: Gary M Weiss, Haym Hirsh
    Abstract:

    Systems that learn from examples often express the learned concept in the form of a disjunctive description. Disjuncts that correctly classify few training examples are known as small Disjuncts and are interesting to machine learning researchers because they have a much higher error rate than large Disjuncts. Previous research has investigated this phenomenon by performing ad hoc analyses of a small number of datasets. In this paper we present a quantitative measure for evaluating the effect of small Disjuncts on learning and use it to analyze 30 benchmark datasets. We investigate the relationship between small Disjuncts and pruning, training set size and noise, and come up with several interesting results.

Margaret Byrne - One of the best experts on this subject based on the ideXlab platform.

  • Variable clonality and genetic structure among disjunct populations of Banksia mimica
    Conservation Genetics, 2020
    Co-Authors: Melissa A. Millar, Margaret Byrne
    Abstract:

    Various factors influence patterns of genetic diversity within and between populations that are important considerations for plant conservation. Both clonality and population genetic differentiation are key factors informing conservation actions, especially for rare species. Banksia mimica is a rare species that occurs in three disjunct locations in the biodiversity hotspot of the southwest Australian Floristic Region of Western Australia. Extant populations are suspected to have varying levels of clonality and high levels of genetic differentiation due to geographic disjunction. A genetic analysis was undertaken in order to confirm clonal reproduction and obtain initial estimates of genotypic diversity, and to assess genetic diversity within, and genetic structure among, populations. Genotypic richness ranged from 0.210 to 1.00 supporting observations in the field of variable degrees of clonal growth. The most clonal populations showed genetic signals of greater levels of observed heterozygosity than expected heterozygosity and negative F_IS values that were not present in other populations of B. mimica or populations of the nonclonal sister taxon B. vestita . There was strong genetic structure with high genetic divergence among geographically disjunct population groups (global F_ST = 0.392, D_ST = 0.475), as is often found within the Australian flora. Genetic differentiation among disjunct populations located on the Whicher Scarp and more northern populations approached, or was greater than, that between Whicher populations and populations of the sister taxa B. vestita . This result is consistent with several other species that show genetic differentiation in disjunct populations located on the Whicher Scarp geomorphological formation. Results suggest a reassessment of the taxonomy and identification of evolutionary significant units for populations of B. mimica would support effective conservation management of this species.

  • extensive long distance pollen dispersal and highly outcrossed mating in historically small and disjunct populations of acacia woodmaniorum fabaceae a rare banded iron formation endemic
    Annals of Botany, 2014
    Co-Authors: Melissa A. Millar, David J Coates, Margaret Byrne
    Abstract:

    BACKGROUND AND AIMS: Understanding patterns of pollen dispersal and variation in mating systems provides insights into the evolutionary potential of plant species and how historically rare species with small disjunct populations persist over long time frames. This study aims to quantify the role of pollen dispersal and the mating system in maintaining contemporary levels of connectivity and facilitating persistence of small populations of the historically rare Acacia woodmaniorum. METHODS: Progeny arrays of A. woodmaniorum were genotyped with nine polymorphic microsatellite markers. A low number of fathers contributed to seed within single pods; therefore, sampling to remove bias of correlated paternity was implemented for further analysis. Pollen immigration and mating system parameters were then assessed in eight populations of varying size and degree of isolation. KEY RESULTS: Pollen immigration into small disjunct populations was extensive (mean minimum estimate 40 % and mean maximum estimate 57 % of progeny) and dispersal occurred over large distances (≤1870m). Pollen immigration resulted in large effective population sizes and was sufficient to ensure adaptive and inbreeding connectivity in small disjunct populations. High outcrossing (mean tm = 0·975) and a lack of apparent inbreeding suggested that a self-incompatibility mechanism is operating. Population parameters, including size and degree of geographic disjunction, were not useful predictors of pollen dispersal or components of the mating system. CONCLUSIONS: Extensive long-distance pollen dispersal and a highly outcrossed mating system are likely to play a key role in maintaining genetic diversity and limiting negative genetic effects of inbreeding and drift in small disjunct populations of A. woodmaniorum. It is proposed that maintenance of genetic connectivity through habitat and pollinator conservation will be a key factor in the persistence of this and other historically rare species with similar extensive long-distance pollen dispersal and highly outcrossed mating systems.