The Experts below are selected from a list of 10398 Experts worldwide ranked by ideXlab platform
Vangipuram Radhakrishna - One of the best experts on this subject based on the ideXlab platform.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
a novel fuzzy gaussian based Dissimilarity Measure for discovering similarity temporal association patterns
Soft Computing, 2018Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, Kimkwang Raymond ChooAbstract:Mining temporal association patterns from time-stamped temporal databases, first introduced in 2009, remain an active area of research. A pattern is temporally similar when it satisfies certain specified subset constraints. The naive and apriori algorithm designed for non-temporal databases cannot be extended to find similar temporal patterns in the context of temporal databases. The brute force approach requires performing \(2^{n }\) true support computations for ‘n’ items; hence, an NP-class problem. Also, the apriori or fp-tree-based algorithms designed for static databases are not directly extendable to temporal databases to retrieve temporal patterns similar to a reference prevalence of user interest. This is because the support of patterns violates the monotonicity property in temporal databases. In our case, support is a vector of values and not a single value. In this paper, we present a novel approach to retrieve temporal association patterns whose prevalence values are similar to those of the user specified reference. This allows us to significantly reduce support computations by defining novel expressions to estimate support bounds. The proposed approach eliminates computational overhead in finding similar temporal patterns. We then introduce a novel Dissimilarity Measure, which is the fuzzy Gaussian-based Dissimilarity Measure. The Measure also holds the monotonicity property. Our evaluations demonstrate that the proposed method outperforms brute force and sequential approaches. We also compare the performance of the proposed approach with the SPAMINE which uses the Euclidean Measure. The proposed approach uses monotonicity property to prune temporal patterns without computing unnecessary true supports and distances.
-
looking into the possibility for designing normal distribution based Dissimilarity Measure to discover time profiled association patterns
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, V Janaki, P V KumarAbstract:This research addresses the design of a novel Dissimilarity Measure for mining similar patterns from time stamped temporal databases applying the concept of standard score and normal distribution. The basic idea behind the design of Dissimilarity Measure is to use and transform supports to z-space and compute the probability of z-score of temporal patterns. The probability is obtained using normal distribution chart. The objective has been to design a normal distribution based Dissimilarity Measure which can be used to discover all valid similarity-profiled temporal association patterns.
-
extending the gaussian membership function for finding similarity between temporal patterns
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Shadi Aljawarneh, Vangipuram Radhakrishna, Aravind CheruvuAbstract:In this paper, the basic Gaussian membership function is extended to design the Dissimilarity Measure. We extend the Dissimilarity Measure proposed in the G-spamine by applying normal distribution. The Dissimilarity Measure proposed in this paper is designed by using the concept of standard normal distribution. For a pattern to be similar, the Dissimilarity between reference and the temporal pattern has to be less than or equal to the Dissimilarity constraint. This Dissimilarity constraint is obtained by transforming the user threshold value to z-space. The Dissimilarity Measure has also been extended to compute the distance bounds by devising necessary expressions. These distance bounds can be used to prune the invalid temporal associations. The algorithm to obtain the similar temporal associations is outlined.
V Janaki - One of the best experts on this subject based on the ideXlab platform.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
looking into the possibility for designing normal distribution based Dissimilarity Measure to discover time profiled association patterns
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, V Janaki, P V KumarAbstract:This research addresses the design of a novel Dissimilarity Measure for mining similar patterns from time stamped temporal databases applying the concept of standard score and normal distribution. The basic idea behind the design of Dissimilarity Measure is to use and transform supports to z-space and compute the probability of z-score of temporal patterns. The probability is obtained using normal distribution chart. The objective has been to design a normal distribution based Dissimilarity Measure which can be used to discover all valid similarity-profiled temporal association patterns.
-
design and analysis of a novel temporal Dissimilarity Measure using gaussian membership function
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Earlier research works addressing the problem of mining time profiled temporal association patterns did not address the possibility of using new similarity Measures in the context of time stamped temporal databases except some of our previous works. This research throws focus on designing a new similarity Measure for mining similarity profiled temporal association patterns. The objective is to design a fuzzy similarity Measure which can be used to discover all valid similarity profiled temporal association patterns.
-
design and analysis of a novel temporal Dissimilarity Measure using gaussian membership function
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Earlier research works addressing the problem of mining time profiled temporal association patterns did not address the possibility of using new similarity Measures in the context of time stamped temporal databases except some of our previous works. This research throws focus on designing a new similarity Measure for mining similarity profiled temporal association patterns. The objective is to design a fuzzy similarity Measure which can be used to discover all valid similarity profiled temporal association patterns.
P V Kumar - One of the best experts on this subject based on the ideXlab platform.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
a novel fuzzy gaussian based Dissimilarity Measure for discovering similarity temporal association patterns
Soft Computing, 2018Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, Kimkwang Raymond ChooAbstract:Mining temporal association patterns from time-stamped temporal databases, first introduced in 2009, remain an active area of research. A pattern is temporally similar when it satisfies certain specified subset constraints. The naive and apriori algorithm designed for non-temporal databases cannot be extended to find similar temporal patterns in the context of temporal databases. The brute force approach requires performing \(2^{n }\) true support computations for ‘n’ items; hence, an NP-class problem. Also, the apriori or fp-tree-based algorithms designed for static databases are not directly extendable to temporal databases to retrieve temporal patterns similar to a reference prevalence of user interest. This is because the support of patterns violates the monotonicity property in temporal databases. In our case, support is a vector of values and not a single value. In this paper, we present a novel approach to retrieve temporal association patterns whose prevalence values are similar to those of the user specified reference. This allows us to significantly reduce support computations by defining novel expressions to estimate support bounds. The proposed approach eliminates computational overhead in finding similar temporal patterns. We then introduce a novel Dissimilarity Measure, which is the fuzzy Gaussian-based Dissimilarity Measure. The Measure also holds the monotonicity property. Our evaluations demonstrate that the proposed method outperforms brute force and sequential approaches. We also compare the performance of the proposed approach with the SPAMINE which uses the Euclidean Measure. The proposed approach uses monotonicity property to prune temporal patterns without computing unnecessary true supports and distances.
-
looking into the possibility for designing normal distribution based Dissimilarity Measure to discover time profiled association patterns
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, V Janaki, P V KumarAbstract:This research addresses the design of a novel Dissimilarity Measure for mining similar patterns from time stamped temporal databases applying the concept of standard score and normal distribution. The basic idea behind the design of Dissimilarity Measure is to use and transform supports to z-space and compute the probability of z-score of temporal patterns. The probability is obtained using normal distribution chart. The objective has been to design a normal distribution based Dissimilarity Measure which can be used to discover all valid similarity-profiled temporal association patterns.
-
design and analysis of a novel temporal Dissimilarity Measure using gaussian membership function
2017 International Conference on Engineering & MIS (ICEMIS), 2017Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Earlier research works addressing the problem of mining time profiled temporal association patterns did not address the possibility of using new similarity Measures in the context of time stamped temporal databases except some of our previous works. This research throws focus on designing a new similarity Measure for mining similarity profiled temporal association patterns. The objective is to design a fuzzy similarity Measure which can be used to discover all valid similarity profiled temporal association patterns.
Shadi Aljawarneh - One of the best experts on this subject based on the ideXlab platform.
-
garuda gaussian Dissimilarity Measure for feature representation and anomaly detection in internet of things
The Journal of Supercomputing, 2020Co-Authors: Shadi Aljawarneh, Radhakrishna VangipuramAbstract:The objective of any anomaly detection system is to efficiently detect several types of malicious traffic patterns that cannot be detected by conventional firewall systems. Designing an efficient intrusion detection system has three primary challenges that include addressing high dimensionality problem, choice of learning algorithm, and distance or similarity Measure used to find the similarity value between any two traffic patterns or input observations. Feature representation and dimensionality reduction have been studied and addressed widely in the literature and have also been applied for the design of intrusion detection systems (IDS). The choice of classifiers is also studied and applied widely in the design of IDS. However, at the heart of IDS lies the choice of distance Measure that is required for an IDS to judge an incoming observation as normal or abnormal. This challenge has been understudied and relatively less addressed in the research literature both from academia and from industry. This research aims at introducing a novel distance Measure that can be used to perform feature clustering and feature representation for efficient intrusion detection. Recent studies such as CANN proposed feature reduction techniques for improving detection and accuracy rates of IDS that used Euclidean distance. However, accuracies of attack classes such as U2R and R2L are not significantly promising. Our approach GARUDA is based on clustering feature patterns incrementally and then representing features in different transformation space through using a novel fuzzy Gaussian Dissimilarity Measure. Experiments are conducted on both KDD and NSL-KDD datasets. The accuracy and detection rates of proposed approach are compared for classifiers such as kNN, J48, naive Bayes, along with CANN and CLAPP approaches. Experiment results proved that proposed approach resulted in the improved accuracy and detection rates for U2R and R2L attack classes when compared to other approaches.
-
GARUDA: Gaussian Dissimilarity Measure for feature representation and anomaly detection in Internet of things
The Journal of Supercomputing, 2020Co-Authors: Shadi Aljawarneh, Radhakrishna VangipuramAbstract:The objective of any anomaly detection system is to efficiently detect several types of malicious traffic patterns that cannot be detected by conventional firewall systems. Designing an efficient intrusion detection system has three primary challenges that include addressing high dimensionality problem, choice of learning algorithm, and distance or similarity Measure used to find the similarity value between any two traffic patterns or input observations. Feature representation and dimensionality reduction have been studied and addressed widely in the literature and have also been applied for the design of intrusion detection systems (IDS). The choice of classifiers is also studied and applied widely in the design of IDS. However, at the heart of IDS lies the choice of distance Measure that is required for an IDS to judge an incoming observation as normal or abnormal. This challenge has been understudied and relatively less addressed in the research literature both from academia and from industry. This research aims at introducing a novel distance Measure that can be used to perform feature clustering and feature representation for efficient intrusion detection. Recent studies such as CANN proposed feature reduction techniques for improving detection and accuracy rates of IDS that used Euclidean distance. However, accuracies of attack classes such as U2R and R2L are not significantly promising. Our approach GARUDA is based on clustering feature patterns incrementally and then representing features in different transformation space through using a novel fuzzy Gaussian Dissimilarity Measure. Experiments are conducted on both KDD and NSL-KDD datasets. The accuracy and detection rates of proposed approach are compared for classifiers such as kNN, J48, naïve Bayes, along with CANN and CLAPP approaches. Experiment results proved that proposed approach resulted in the improved accuracy and detection rates for U2R and R2L attack classes when compared to other approaches.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
astra a novel interest Measure for unearthing latent temporal associations and trends through extending basic gaussian membership function
Multimedia Tools and Applications, 2019Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, V JanakiAbstract:Time profiled association mining is one of the important and challenging research problems that is relatively less addressed. Time profiled association mining has two main challenges that must be addressed. These include addressing i) Dissimilarity Measure that also holds monotonicity property and can efficiently prune itemset associations ii) approaches for estimating prevalence values of itemset associations over time. The pioneering research that addressed time profiled association mining is by J.S. Yoo using Euclidean distance. It is widely known fact that this distance Measure suffers from high dimensionality. Given a time stamped transaction database, time profiled association mining refers to the discovery of underlying and hidden time profiled itemset associations whose true prevalence variations are similar as the user query sequence under subset constraints that include i) allowable Dissimilarity value ii) a reference query time sequence iii) Dissimilarity function that can find degree of similarity between a temporal itemset and reference. In this paper, we propose a novel Dissimilarity Measure whose design is a function of product based gaussian membership function through extending the similarity function proposed in our earlier research (G-Spamine). Our approach, MASTER (Mining of Similar Temporal Associations) which is primarily inspired from SPAMINE uses the Dissimilarity Measure proposed in this paper and support bound estimation approach proposed in our earlier research. Expression for computation of distance bounds of temporal patterns are designed considering the proposed Measure and support estimation approach. Experiments are performed by considering naive, sequential, Spamine and G-Spamine approaches under various test case considerations that study the scalability and computational performance of the proposed approach. Experimental results prove the scalability and efficiency of the proposed approach. The correctness and completeness of proposed approach is also proved analytically.
-
a novel fuzzy gaussian based Dissimilarity Measure for discovering similarity temporal association patterns
Soft Computing, 2018Co-Authors: Vangipuram Radhakrishna, Shadi Aljawarneh, P V Kumar, Kimkwang Raymond ChooAbstract:Mining temporal association patterns from time-stamped temporal databases, first introduced in 2009, remain an active area of research. A pattern is temporally similar when it satisfies certain specified subset constraints. The naive and apriori algorithm designed for non-temporal databases cannot be extended to find similar temporal patterns in the context of temporal databases. The brute force approach requires performing \(2^{n }\) true support computations for ‘n’ items; hence, an NP-class problem. Also, the apriori or fp-tree-based algorithms designed for static databases are not directly extendable to temporal databases to retrieve temporal patterns similar to a reference prevalence of user interest. This is because the support of patterns violates the monotonicity property in temporal databases. In our case, support is a vector of values and not a single value. In this paper, we present a novel approach to retrieve temporal association patterns whose prevalence values are similar to those of the user specified reference. This allows us to significantly reduce support computations by defining novel expressions to estimate support bounds. The proposed approach eliminates computational overhead in finding similar temporal patterns. We then introduce a novel Dissimilarity Measure, which is the fuzzy Gaussian-based Dissimilarity Measure. The Measure also holds the monotonicity property. Our evaluations demonstrate that the proposed method outperforms brute force and sequential approaches. We also compare the performance of the proposed approach with the SPAMINE which uses the Euclidean Measure. The proposed approach uses monotonicity property to prune temporal patterns without computing unnecessary true supports and distances.
Z. Huang - One of the best experts on this subject based on the ideXlab platform.
-
A fuzzy k-modes algorithm for clustering categorical data
IEEE Transactions on Fuzzy Systems, 1999Co-Authors: Z. HuangAbstract:This correspondence describes extensions to the fuzzy k-means algorithm for clustering categorical data. By using a simple matching Dissimilarity Measure for categorical objects and modes instead of means for clusters, a new approach is developed, which allows the use of the k-means paradigm to efficiently cluster large categorical data sets. A fuzzy k-modes algorithm is presented and the effectiveness of the algorithm is demonstrated with experimental results.
-
Extensions to the k-Means Algorithm for Clustering Large Data Sets with Categorical Values
Data Mining and Knowledge Discovery, 1998Co-Authors: Z. HuangAbstract:The k-means algorithm is well known for its efficiency in clustering large data sets. However, working only on numeric values prohibits it from being used to cluster real world data containing categorical values. In this paper we present two algorithms which extend the k-means algorithm to categorical domains and domains with mixed numeric and categorical values. The k-modes algorithm uses a simple matching Dissimilarity Measure to deal with categorical objects, replaces the means of clusters with modes, and uses a frequency-based method to update modes in the clustering process to minimise the clustering cost function. With these extensions the k-modes algorithm enables the clustering of categorical data in a fashion similar to k-means. The k-prototypes algorithm, through the definition of a combined Dissimilarity Measure, further integrates the k-means and k-modes algorithms to allow for clustering objects described by mixed numeric and categorical attributes. We use the well known soybean disease and credit approval data sets to demonstrate the clustering performance of the two algorithms. Our experiments on two real world data sets with half a million objects each show that the two algorithms are efficient when clustering large data sets, which is critical to data mining applications.