The Experts below are selected from a list of 18513 Experts worldwide ranked by ideXlab platform
Yihong Gong - One of the best experts on this subject based on the ideXlab platform.
-
Integrating Document Clustering and MultiDocument Summarization
ACM Transactions on Knowledge Discovery from Data, 2011Co-Authors: Dingding Wang, Shenghuo Zhu, Yun Chi, Yihong GongAbstract:Document understanding techniques such as Document Clustering and multiDocument summarization have been receiving much attention recently. Current Document Clustering methods usually represent the given collection of Documents as a Document-term matrix and then conduct the Clustering process. Although many of these Clustering methods can group the Documents effectively, it is still hard for people to capture the meaning of the Documents since there is no satisfactory interpretation for each Document cluster. A straightforward solution is to first cluster the Documents and then summarize each Document cluster using summarization methods. However, most of the current summarization methods are solely based on the sentence-term matrix and ignore the context dependence of the sentences. As a result, the generated summaries lack guidance from the Document clusters. In this article, we propose a new language model to simultaneously cluster and summarize Documents by making use of both the Document-term and sentence-term matrices. By utilizing the mutual influence of Document Clustering and summarization, our method makes; (1) a better Document Clustering method with more meaningful interpretation; and (2) an effective Document summarization method with guidance from Document Clustering. Experimental results on various Document datasets show the effectiveness of our proposed method and the high interpretability of the generated summaries.
-
Document Clustering based on non negative matrix factorization
International ACM SIGIR Conference on Research and Development in Information Retrieval, 2003Co-Authors: Xin Liu, Yihong GongAbstract:In this paper, we propose a novel Document Clustering method based on the non-negative factorization of the term-Document matrix of the given Document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular Document cluster, and each Document is represented as an additive combination of the base topics. The cluster membership of each Document can be easily determined by finding the base topic (the axis) with which the Document has the largest projection value. Our experimental evaluations show that the proposed Document Clustering method surpasses the latent semantic indexing and the spectral Clustering methods not only in the easy and reliable derivation of Document Clustering results, but also in Document Clustering accuracies.
-
SIGIR - Document Clustering based on non-negative matrix factorization
Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval - SIGIR '03, 2003Co-Authors: Xin Liu, Yihong GongAbstract:In this paper, we propose a novel Document Clustering method based on the non-negative factorization of the term-Document matrix of the given Document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular Document cluster, and each Document is represented as an additive combination of the base topics. The cluster membership of each Document can be easily determined by finding the base topic (the axis) with which the Document has the largest projection value. Our experimental evaluations show that the proposed Document Clustering method surpasses the latent semantic indexing and the spectral Clustering methods not only in the easy and reliable derivation of Document Clustering results, but also in Document Clustering accuracies.
-
Document Clustering with cluster refinement and model selection capabilities
International ACM SIGIR Conference on Research and Development in Information Retrieval, 2002Co-Authors: Xin Liu, Yihong Gong, Shenghuo ZhuAbstract:In this paper, we propose a Document Clustering method that strives to achieve: (1) a high accuracy of Document Clustering, and (2) the capability of estimating the number of clusters in the Document corpus (i.e. the model selection capability). To accurately cluster the given Document corpus, we employ a richer feature set to represent each Document, and use the Gaussian Mixture Model (GMM) together with the Expectation-Maximization (EM) algorithm to conduct an initial Document Clustering. From this initial result, we identify a set of discriminative featuresfor each cluster, and refine the initially obtained Document clusters by voting on the cluster label of each Document using this discriminative feature set. This self-refinement process of discriminative feature identification and cluster label voting is iteratively applied until the convergence of Document clusters. On the other hand, the model selection capability is achieved by introducing randomness in the cluster initialization stage, and then discovering a value C for the number of clusters N by which running the Document Clustering process for a fixed number of times yields sufficiently similar results. Performance evaluations exhibit clear superiority of the proposed method with its improved Document Clustering and model selection accuracies. The evaluations also demonstrate how each feature as well as the cluster refinement process contribute to the Document Clustering accuracy.
Xin Liu - One of the best experts on this subject based on the ideXlab platform.
-
Document Clustering based on non negative matrix factorization
International ACM SIGIR Conference on Research and Development in Information Retrieval, 2003Co-Authors: Xin Liu, Yihong GongAbstract:In this paper, we propose a novel Document Clustering method based on the non-negative factorization of the term-Document matrix of the given Document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular Document cluster, and each Document is represented as an additive combination of the base topics. The cluster membership of each Document can be easily determined by finding the base topic (the axis) with which the Document has the largest projection value. Our experimental evaluations show that the proposed Document Clustering method surpasses the latent semantic indexing and the spectral Clustering methods not only in the easy and reliable derivation of Document Clustering results, but also in Document Clustering accuracies.
-
SIGIR - Document Clustering based on non-negative matrix factorization
Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval - SIGIR '03, 2003Co-Authors: Xin Liu, Yihong GongAbstract:In this paper, we propose a novel Document Clustering method based on the non-negative factorization of the term-Document matrix of the given Document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular Document cluster, and each Document is represented as an additive combination of the base topics. The cluster membership of each Document can be easily determined by finding the base topic (the axis) with which the Document has the largest projection value. Our experimental evaluations show that the proposed Document Clustering method surpasses the latent semantic indexing and the spectral Clustering methods not only in the easy and reliable derivation of Document Clustering results, but also in Document Clustering accuracies.
-
Document Clustering with cluster refinement and model selection capabilities
International ACM SIGIR Conference on Research and Development in Information Retrieval, 2002Co-Authors: Xin Liu, Yihong Gong, Shenghuo ZhuAbstract:In this paper, we propose a Document Clustering method that strives to achieve: (1) a high accuracy of Document Clustering, and (2) the capability of estimating the number of clusters in the Document corpus (i.e. the model selection capability). To accurately cluster the given Document corpus, we employ a richer feature set to represent each Document, and use the Gaussian Mixture Model (GMM) together with the Expectation-Maximization (EM) algorithm to conduct an initial Document Clustering. From this initial result, we identify a set of discriminative featuresfor each cluster, and refine the initially obtained Document clusters by voting on the cluster label of each Document using this discriminative feature set. This self-refinement process of discriminative feature identification and cluster label voting is iteratively applied until the convergence of Document clusters. On the other hand, the model selection capability is achieved by introducing randomness in the cluster initialization stage, and then discovering a value C for the number of clusters N by which running the Document Clustering process for a fixed number of times yields sufficiently similar results. Performance evaluations exhibit clear superiority of the proposed method with its improved Document Clustering and model selection accuracies. The evaluations also demonstrate how each feature as well as the cluster refinement process contribute to the Document Clustering accuracy.
Ruizhang Huang - One of the best experts on this subject based on the ideXlab platform.
-
Adjusting the Inheritance of Topic for Dynamic Document Clustering
Communications in Computer and Information Science, 2019Co-Authors: Ruizhang Huang, Yingxue Zhu, Yanping Chen, Yue Yang, Yang Jian, Yaru MengAbstract:Organizing streaming Documents from time-varying dataset is meaningful but difficult because topics evolve over time. Dynamic Document Clustering is a vital research problem, which helps to group the time-varying Documents into a number of clusters corresponding to their underlying topics. Datasets are partitioned into a set of time slides to transfer the streaming Document Clustering from a continuous problem to a categorical one. Traditional dynamic Document Clustering approach tends to inherit topic information over time directly with no consideration of the nature of datasets. In this paper, we design a novel prior-adjusted dynamic Document Clustering approach, namely PADC, which is able to adjust the topic inheritance process according to two important of datasets characteristics, in particular, the interval between dataset time slides and the size of dataset time slides. A collapsed Gibbs sampling algorithm is investigated to infer the Document structure for all time slides with underlying time-varying topics. Parameters for underlying topics inheritance, as well as parameters of the number of clusters in each time slide, are estimated simultaneously. Extensive experiments have been conducted comparing the PADC model with state-of-the-art dynamic Document Clustering approaches. Experimental results demonstrate that the PADC model is robust and effective for the dynamic Document Clustering problem.
-
Dirichlet Process Mixture Model for Document Clustering with Feature Partition
IEEE Transactions on Knowledge and Data Engineering, 2013Co-Authors: Ruizhang Huang, Zhaojun Wang, Jun Zhang, Liangxing ShiAbstract:Finding the appropriate number of clusters to which Documents should be partitioned is crucial in Document Clustering. In this paper, we propose a novel approach, namely DPMFP, to discover the latent cluster structure based on the DPM model without requiring the number of clusters as input. Document features are automatically partitioned into two groups, in particular, discriminative words and nondiscriminative words, and contribute differently to Document Clustering. A variational inference algorithm is investigated to infer the Document collection structure as well as the partition of Document words at the same time. Our experiments indicate that our proposed approach performs well on the synthetic data set as well as real data sets. The comparison between our approach and state-of-the-art Document Clustering approaches shows that our approach is robust and effective for Document Clustering.
-
Semi-supervised Document Clustering with active learning
2008Co-Authors: Wai Lam, Ruizhang HuangAbstract:This thesis presents a new framework for automatically partitioning text Documents taking into consideration of constraints given by users. Semi-supervised Document Clustering is developed based on pairwise constraints. Different from traditional semi-supervised Document Clustering approaches which assume pairwise constraints to be prepared by user beforehand, we develop a novel framework for automatically discovering pairwise constraints revealing the user grouping preference. Active learning approach for choosing informative Document pairs is designed by measuring the amount of information that can be obtained by revealing judgments of Document pairs. For this purpose, three models, namely, uncertainty model, generation error model, and term-to-term relationship model, are designed for measuring the informativeness of Document pairs from different perspectives. Dependent active learning approach is developed by extending the active learning approach to avoid redundant Document pair selection. Two models are investigated for estimating the likelihood that a Document pair is redundant to previously selected Document pairs, namely, KL divergence model and symmetric model. Most existing semi-supervised Document Clustering approaches are model-based Clustering and can be treated as parametric model taking an assumption that the underlying clusters follow a certain pre-defined distribution. In our semi-supervised Document Clustering, each cluster is represented by a non-parametric probability distribution. Two approaches are designed for incorporating pairwise constraints in the Document Clustering approach. The first approach, term-to-term relationship approach (TR), uses pairwise constraints for capturing term-to-term dependence relationships. The second approach, linear combination approach (LC), combines the Clustering objective function with the user-provided constraints linearly. Extensive experimental results show that our proposed framework is effective.
Han-wei Hsiao - One of the best experts on this subject based on the ideXlab platform.
-
A collaborative filtering-based approach to personalized Document Clustering
Decision Support Systems, 2008Co-Authors: Chih-ping Wei, Chin-sheng Yang, Han-wei HsiaoAbstract:Document Clustering is an intentional act that reflects individual preferences with regard to the semantic coherency and relevant categorization of Documents. Hence, effective Document Clustering must consider individual preferences and needs to support personalization in Document categorization. Most existing Document-Clustering techniques, generally anchoring in pure content-based analysis, generate a single set of clusters for all individuals without tailoring to individuals' preferences and thus are unable to support personalization. The partial-Clustering-based personalized Document-Clustering approach, incorporating a target individual's partial Clustering into the Document-Clustering process, has been proposed to facilitate personalized Document Clustering. However, given a collection of Documents to be clustered, the individual might have categorized only a small subset of the collection into his or her personal folders. In this case, the small partial Clustering would degrade the effectiveness of the existing personalized Document-Clustering approach for this particular individual. In response, we extend this approach and propose the collaborative-filtering-based personalized Document-Clustering (CFC) technique that expands the size of an individual's partial Clustering by considering those of other users with similar categorization preferences. Our empirical evaluation results suggest that when given a small-sized partial Clustering established by an individual, the proposed CFC technique generally achieves better Clustering effectiveness for the individual than does the partial-Clustering-based personalized Document-Clustering technique.
-
PACIS - Personalized Document Clustering: A Collaborative-Filtering-Based Approach
2004Co-Authors: Chih-ping Wei, Chin-sheng Yang, Han-wei HsiaoAbstract:To manage the ever-increasing volume of Documents, individuals and organizations frequently organize their Documents into categories that facilitate Document management and subsequent information access and browsing. However, Document Clustering is intentional acts that reflect individual preferences with regard to the semantic coherency and relevant categorization of Documents. Hence, an effective Document Clustering must consider individual preferences and needs to support personalization in Document categorization. In this study, we design and implement a collaborative-filtering-based Document-Clustering (CFC) technique by incorporating an individual’s and his/her neighbors’ partial Clusterings for supporting personalized Document Clustering. The empirical evaluation results suggest that the use of an individual’s partial Clustering can achieve a better personalized Clustering result than does the content-based Document Clustering technique. Moreover, use of the collaborative-filtering approach for expanding an individual’s partial Clustering can further improve personalized Clustering, measured by cluster recall and precision.
Chih-ping Wei - One of the best experts on this subject based on the ideXlab platform.
-
A collaborative filtering-based approach to personalized Document Clustering
Decision Support Systems, 2008Co-Authors: Chih-ping Wei, Chin-sheng Yang, Han-wei HsiaoAbstract:Document Clustering is an intentional act that reflects individual preferences with regard to the semantic coherency and relevant categorization of Documents. Hence, effective Document Clustering must consider individual preferences and needs to support personalization in Document categorization. Most existing Document-Clustering techniques, generally anchoring in pure content-based analysis, generate a single set of clusters for all individuals without tailoring to individuals' preferences and thus are unable to support personalization. The partial-Clustering-based personalized Document-Clustering approach, incorporating a target individual's partial Clustering into the Document-Clustering process, has been proposed to facilitate personalized Document Clustering. However, given a collection of Documents to be clustered, the individual might have categorized only a small subset of the collection into his or her personal folders. In this case, the small partial Clustering would degrade the effectiveness of the existing personalized Document-Clustering approach for this particular individual. In response, we extend this approach and propose the collaborative-filtering-based personalized Document-Clustering (CFC) technique that expands the size of an individual's partial Clustering by considering those of other users with similar categorization preferences. Our empirical evaluation results suggest that when given a small-sized partial Clustering established by an individual, the proposed CFC technique generally achieves better Clustering effectiveness for the individual than does the partial-Clustering-based personalized Document-Clustering technique.
-
PACIS - Context-aware Document-Clustering Technique
2007Co-Authors: Chin-sheng Yang, Chih-ping WeiAbstract:Document Clustering is an intentional act that should reflect individuals’ preferences with regard to the semantic coherency or relevant categorization of Documents and should conform to the context of a target task under investigation. Thus, effective DocumentClustering techniques need to take into account a user’s categorization context defined by or relevant to the target task under consideration. However, existing Document-Clustering techniques generally anchor in pure content-based analysis and therefore are not able to facilitate context-aware Document-Clustering. In response, we propose a Context-Aware Document-Clustering (CAC) technique that takes into consideration a user’s categorization preference (expressed as a list of anchoring terms) relevant to the context of a target task and subsequently generates a set of Document clusters from this specific contextual perspective. Our empirical evaluation results suggest that our proposed CAC technique outperforms the pure content-based Document-Clustering technique.
-
PACIS - Personalized Document Clustering: A Collaborative-Filtering-Based Approach
2004Co-Authors: Chih-ping Wei, Chin-sheng Yang, Han-wei HsiaoAbstract:To manage the ever-increasing volume of Documents, individuals and organizations frequently organize their Documents into categories that facilitate Document management and subsequent information access and browsing. However, Document Clustering is intentional acts that reflect individual preferences with regard to the semantic coherency and relevant categorization of Documents. Hence, an effective Document Clustering must consider individual preferences and needs to support personalization in Document categorization. In this study, we design and implement a collaborative-filtering-based Document-Clustering (CFC) technique by incorporating an individual’s and his/her neighbors’ partial Clusterings for supporting personalized Document Clustering. The empirical evaluation results suggest that the use of an individual’s partial Clustering can achieve a better personalized Clustering result than does the content-based Document Clustering technique. Moreover, use of the collaborative-filtering approach for expanding an individual’s partial Clustering can further improve personalized Clustering, measured by cluster recall and precision.