The Experts below are selected from a list of 21015 Experts worldwide ranked by ideXlab platform

Jia Zhao - One of the best experts on this subject based on the ideXlab platform.

  • a novel process based association rule approach through maximal frequent itemsets for big data processing
    Future Generation Computer Systems, 2018
    Co-Authors: Zelei Liu, Yan Ding, Quangang Wen, Jia Zhao
    Abstract:

    The maximal frequent itemsets issue in big data processing has become a hot research topic. Most of the previous work on big data processing directly analyzes the data through the existing approaches, which would cause problems of redundant computation, high time complexity, and large storage space. To solve the problems, this paper proposes a Heuristic MapReduce-based Association rule approach through Maximal frequent itemsets mining, HMAM. The main idea is: At first, by directly operating on the Transaction Database, we allocate Transactions to different processing nodes and group all Transactions according to dimension. Then, we screen the most frequent Transactions from each Transaction set using the Bitmap-Sort and obtain best-Transaction-set through aggregating all the Transaction-elects of each Transaction set. The current candidate maximal frequent itemsets can be acquired by removing sub-Transactions in terms of the inclusion relations of the items in best-Transaction-set. At the same time, each subset of sub-Transactions in the candidate maximal frequent itemsets is discarded from all Transaction sets. Then the final candidate maximal frequent itemsets can be obtained by iteration until each Transaction set is empty. Finally, we achieve the acquisition of maximal frequent itemsets by employing the minimum support threshold. The experimental results demonstrate that compared with the existing approaches, HMAM significantly avoids producing a large number of candidate itmesets resulting from join operation, accelerates the speed of mining the maximal frequent itemsets, and improves the utilization rate of resources simultaneously. Allocate Transactions to different nodes and group them in terms of dimension.Screen the most frequent Transactions from each Transaction set using Bitmap-Sort.Obtain maximal frequent itemsets by employing the minimum support threshold.

Ming-syan Chen - One of the best experts on this subject based on the ideXlab platform.

  • sliding window filtering an efficient algorithm for incremental mining
    Conference on Information and Knowledge Management, 2001
    Co-Authors: Ming-syan Chen
    Abstract:

    We explore in this paper an effective sliding-window filtering (abbreviatedly as SWF) algorithm for incremental mining of association rules. In essence, by partitioning a Transaction Database into several partitions, algorithm SWF employs a filtering threshold in each partition to deal with the candidate itemset generation. Under SWF, the cumulative information of mining previous partitions is selectively carried over toward the generation of candidate itemsets for the subsequent partitions. Algorithm SWF not only significantly reduces I/O and CPU cost by the concepts of cumulative filtering and scan reduction techniques but also effectively controls memory utilization by the technique of sliding-window partition. Algorithm SWF is particularly powerful for efficient incremental mining for an ongoing time-variant Transaction Database. By utilizing proper scan reduction techniques, only one scan of the incremented dataset is needed by algorithm SWF. The I/O cost of SWF is, in orders of magnitude, smaller than those required by prior methods, thus resolving the performance bottleneck. Experimental studies are performed to evaluate performance of algorithm SWF. It is noted that the improvement achieved by algorithm SWF is even more prominent as the incremented portion of the dataset increases and also as the size of the Database increases.

  • Using a hash-based method with Transaction trimming for mining association rules
    IEEE Transactions on Knowledge and Data Engineering, 1997
    Co-Authors: Jong Soo Park, Ming-syan Chen
    Abstract:

    We examine the issue of mining association rules among items in a large Database of sales Transactions. Mining association rules means that, given a Database of sales Transactions, to discover all associations among items such that the presence of some items in a Transaction will imply the presence of other items in the same Transaction. The mining of association rules can be mapped into the problem of discovering large itemsets where a large itemset is a group of items that appear in a sufficient number of Transactions. The problem of discovering large itemsets can be solved by constructing a candidate set of itemsets first, and then, identifying, within this candidate set, these itemsets that meet the large itemset requirement. Generally, this is done iteratively for each large k-itemset in increasing order of k, where a large k-itemset is a large itemset with k items. To determine large itemsets from a huge number of candidate sets in early iterations is usually the dominating factor for the overall data mining performance. To address this issue, we develop an effective algorithm for the candidate set generation. It is a hash-based algorithm and is especially effective for the generation of a candidate set for large 2-itemsets. Explicitly, the number of candidate 2-itemsets generated by the proposed algorithm is, in orders of magnitude, smaller than that by previous methods, thus resolving the performance bottleneck. Note that the generation of smaller candidate sets enables us to effectively trim the Transaction Database size at a much earlier stage of the iterations, thereby reducing the computational cost for later iterations significantly. The advantage of the proposed algorithm also provides us the opportunity of reducing the amount of disk I/O required. An extensive simulation study is conducted to evaluate performance of the proposed algorithm.

  • an effective hash based algorithm for mining association rules
    International Conference on Management of Data, 1995
    Co-Authors: Jong Soo Park, Ming-syan Chen, Philip S Yu
    Abstract:

    In this paper, we examine the issue of mining association rules among items in a large Database of sales Transactions. The mining of association rules can be mapped into the problem of discovering large itemsets where a large itemset is a group of items which appear in a sufficient number of Transactions. The problem of discovering large itemsets can be solved by constructing a candidate set of itemsets first and then, identifying, within this candidate set, those itemsets that meet the large itemset requirement. Generally this is done iteratively for each large k-itemset in increasing order of k where a large k-itemset is a large itemset with k items. To determine large itemsets from a huge number of candidate large itemsets in early iterations is usually the dominating factor for the overall data mining performance. To address this issue, we propose an effective hash-based algorithm for the candidate set generation. Explicitly, the number of candidate 2-itemsets generated by the proposed algorithm is, in orders of magnitude, smaller than that by previous methods, thus resolving the performance bottleneck. Note that the generation of smaller candidate sets enables us to effectively trim the Transaction Database size at a much earlier stage of the iterations, thereby reducing the computational cost for later iterations significantly. Extensive simulation study is conducted to evaluate performance of the proposed algorithm.

Ajith Abraham - One of the best experts on this subject based on the ideXlab platform.

  • an efficient algorithm for incremental mining of temporal association rules
    Data and Knowledge Engineering, 2010
    Co-Authors: Tarek F Gharib, Hamed Nassar, Mohamed Taha, Ajith Abraham
    Abstract:

    This paper presents the concept of temporal association rules in order to solve the problem of handling time series by including time expressions into association rules. Actually, temporal Databases are continually appended or updated so that the discovered rules need to be updated. Re-running the temporal mining algorithm every time is ineffective since it neglects the previously discovered rules, and repeats the work done previously. Furthermore, existing incremental mining techniques cannot deal with temporal association rules. In this paper, an incremental algorithm to maintain the temporal association rules in a Transaction Database is proposed. The algorithm benefits from the results of earlier mining to derive the final mining output. The experimental results on both the synthetic and the real dataset illustrate a significant improvement over the conventional approach of mining the entire updated Database.

Philip S Yu - One of the best experts on this subject based on the ideXlab platform.

  • hitfraud a broad learning approach for collective fraud detection in heterogeneous information networks
    International Conference on Data Mining, 2017
    Co-Authors: Siim Viidu, Philip S Yu
    Abstract:

    On electronic game platforms, different payment Transactions have different levels of risk. Risk is generally higher for digital goods in e-commerce. However, it differs based on product and its popularity, the offer type (packaged game, virtual currency to a game or subscription service), storefront and geography. Existing fraud policies and models make decisions independently for each Transaction based on Transaction attributes, payment velocities, user characteristics, and other relevant information. However, suspicious Transactions may still evade detection and hence we propose a broad learning approach leveraging a graph based perspective to uncover relationships among suspicious Transactions, i.e., inter-Transaction dependency. Our focus is to detect suspicious Transactions by capturing common fraudulent behaviors that would not be considered suspicious when being considered in isolation. In this paper, we present HitFraud that leverages heterogeneous information networks for collective fraud detection by exploring correlated and fast evolving fraudulent behaviors. First, a heterogeneous information network is designed to link entities of interest in the Transaction Database via different semantics. Then, graph based features are efficiently discovered from the network exploiting the concept of meta-paths, and decisions on frauds are made collectively on test instances. Experiments on real-world payment Transaction data from Electronic Arts demonstrate that the prediction performance is effectively boosted by HitFraud with fast convergence.

  • hitfraud a broad learning approach for collective fraud detection in heterogeneous information networks
    arXiv: Learning, 2017
    Co-Authors: Siim Viidu, Philip S Yu
    Abstract:

    On electronic game platforms, different payment Transactions have different levels of risk. Risk is generally higher for digital goods in e-commerce. However, it differs based on product and its popularity, the offer type (packaged game, virtual currency to a game or subscription service), storefront and geography. Existing fraud policies and models make decisions independently for each Transaction based on Transaction attributes, payment velocities, user characteristics, and other relevant information. However, suspicious Transactions may still evade detection and hence we propose a broad learning approach leveraging a graph based perspective to uncover relationships among suspicious Transactions, i.e., inter-Transaction dependency. Our focus is to detect suspicious Transactions by capturing common fraudulent behaviors that would not be considered suspicious when being considered in isolation. In this paper, we present HitFraud that leverages heterogeneous information networks for collective fraud detection by exploring correlated and fast evolving fraudulent behaviors. First, a heterogeneous information network is designed to link entities of interest in the Transaction Database via different semantics. Then, graph based features are efficiently discovered from the network exploiting the concept of meta-paths, and decisions on frauds are made collectively on test instances. Experiments on real-world payment Transaction data from Electronic Arts demonstrate that the prediction performance is effectively boosted by HitFraud with fast convergence where the computation of meta-path based features is largely optimized. Notably, recall can be improved up to 7.93% and F-score 4.62% compared to baselines.

  • an effective hash based algorithm for mining association rules
    International Conference on Management of Data, 1995
    Co-Authors: Jong Soo Park, Ming-syan Chen, Philip S Yu
    Abstract:

    In this paper, we examine the issue of mining association rules among items in a large Database of sales Transactions. The mining of association rules can be mapped into the problem of discovering large itemsets where a large itemset is a group of items which appear in a sufficient number of Transactions. The problem of discovering large itemsets can be solved by constructing a candidate set of itemsets first and then, identifying, within this candidate set, those itemsets that meet the large itemset requirement. Generally this is done iteratively for each large k-itemset in increasing order of k where a large k-itemset is a large itemset with k items. To determine large itemsets from a huge number of candidate large itemsets in early iterations is usually the dominating factor for the overall data mining performance. To address this issue, we propose an effective hash-based algorithm for the candidate set generation. Explicitly, the number of candidate 2-itemsets generated by the proposed algorithm is, in orders of magnitude, smaller than that by previous methods, thus resolving the performance bottleneck. Note that the generation of smaller candidate sets enables us to effectively trim the Transaction Database size at a much earlier stage of the iterations, thereby reducing the computational cost for later iterations significantly. Extensive simulation study is conducted to evaluate performance of the proposed algorithm.

Zelei Liu - One of the best experts on this subject based on the ideXlab platform.

  • a novel process based association rule approach through maximal frequent itemsets for big data processing
    Future Generation Computer Systems, 2018
    Co-Authors: Zelei Liu, Yan Ding, Quangang Wen, Jia Zhao
    Abstract:

    The maximal frequent itemsets issue in big data processing has become a hot research topic. Most of the previous work on big data processing directly analyzes the data through the existing approaches, which would cause problems of redundant computation, high time complexity, and large storage space. To solve the problems, this paper proposes a Heuristic MapReduce-based Association rule approach through Maximal frequent itemsets mining, HMAM. The main idea is: At first, by directly operating on the Transaction Database, we allocate Transactions to different processing nodes and group all Transactions according to dimension. Then, we screen the most frequent Transactions from each Transaction set using the Bitmap-Sort and obtain best-Transaction-set through aggregating all the Transaction-elects of each Transaction set. The current candidate maximal frequent itemsets can be acquired by removing sub-Transactions in terms of the inclusion relations of the items in best-Transaction-set. At the same time, each subset of sub-Transactions in the candidate maximal frequent itemsets is discarded from all Transaction sets. Then the final candidate maximal frequent itemsets can be obtained by iteration until each Transaction set is empty. Finally, we achieve the acquisition of maximal frequent itemsets by employing the minimum support threshold. The experimental results demonstrate that compared with the existing approaches, HMAM significantly avoids producing a large number of candidate itmesets resulting from join operation, accelerates the speed of mining the maximal frequent itemsets, and improves the utilization rate of resources simultaneously. Allocate Transactions to different nodes and group them in terms of dimension.Screen the most frequent Transactions from each Transaction set using Bitmap-Sort.Obtain maximal frequent itemsets by employing the minimum support threshold.