The Experts below are selected from a list of 12708 Experts worldwide ranked by ideXlab platform

Eren Manavoglu - One of the best experts on this subject based on the ideXlab platform.

  • rule based word clustering for Document Metadata extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Kostas Tsioutsiouliklis, Lee C Giles, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

  • SAC - Rule-based word clustering for Document Metadata extraction
    Proceedings of the 2005 ACM symposium on Applied computing - SAC '05, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, C. Lee Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

  • automatic Document Metadata extraction using support vector machines
    ACM IEEE Joint Conference on Digital Libraries, 2003
    Co-Authors: C L Giles, Eren Manavoglu, Zhenyue Zhang
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata extraction. We describe a support vector machine classification-based method for Metadata extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata extraction method was originally designed to improve the Metadata extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

  • JCDL - Automatic Document Metadata extraction using support vector machines
    2003 Joint Conference on Digital Libraries 2003. Proceedings., 2003
    Co-Authors: Hui Han, C L Giles, Eren Manavoglu, Zhenyue Zhang, Hongyuan Zha, Edward A. Fox
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata extraction. We describe a support vector machine classification-based method for Metadata extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata extraction method was originally designed to improve the Metadata extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

Hui Han - One of the best experts on this subject based on the ideXlab platform.

  • SAC - Rule-based word clustering for Document Metadata extraction
    Proceedings of the 2005 ACM symposium on Applied computing - SAC '05, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, C. Lee Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

  • rule based word clustering for Document Metadata extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Kostas Tsioutsiouliklis, Lee C Giles, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

  • JCDL - Automatic Document Metadata extraction using support vector machines
    2003 Joint Conference on Digital Libraries 2003. Proceedings., 2003
    Co-Authors: Hui Han, C L Giles, Eren Manavoglu, Zhenyue Zhang, Hongyuan Zha, Edward A. Fox
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata extraction. We describe a support vector machine classification-based method for Metadata extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata extraction method was originally designed to improve the Metadata extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

Xiangmin Zhang - One of the best experts on this subject based on the ideXlab platform.

  • rule based word clustering for Document Metadata extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Kostas Tsioutsiouliklis, Lee C Giles, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

  • SAC - Rule-based word clustering for Document Metadata extraction
    Proceedings of the 2005 ACM symposium on Applied computing - SAC '05, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, C. Lee Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science Document search engine, uses automatic text classification methods for Document indexing. Text classification uses a Document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of Document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of Document Metadata extraction.

Zhenyue Zhang - One of the best experts on this subject based on the ideXlab platform.

  • automatic Document Metadata extraction using support vector machines
    ACM IEEE Joint Conference on Digital Libraries, 2003
    Co-Authors: C L Giles, Eren Manavoglu, Zhenyue Zhang
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata extraction. We describe a support vector machine classification-based method for Metadata extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata extraction method was originally designed to improve the Metadata extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

  • JCDL - Automatic Document Metadata extraction using support vector machines
    2003 Joint Conference on Digital Libraries 2003. Proceedings., 2003
    Co-Authors: Hui Han, C L Giles, Eren Manavoglu, Zhenyue Zhang, Hongyuan Zha, Edward A. Fox
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata extraction. We describe a support vector machine classification-based method for Metadata extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata extraction method was originally designed to improve the Metadata extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

Birger Larsen - One of the best experts on this subject based on the ideXlab platform.

  • CLEF (Online Working Notes/Labs/Workshop) - RSLIS at INEX 2012: Social Book Search Track
    2012
    Co-Authors: Toine Bogers, Birger Larsen
    Abstract:

    In this paper, we describe our participation in the INEX 2012 So- cial Book Search track. We investigate the contribution of different types of Document Metadata, both social and controlled, and examine the effective- ness of re-ranking retrieval results using different social features, such as user ratings, tags, and authorship information. We find that the best results are obtained using all available Document fields and topic representations. Re- ranking retrieval results works better on shorter topic representations, where there is less information for the retrieval algorithm to work with; longer topic representations do not benefit from our social re-ranking approaches.

  • INEX - RSLIS at INEX 2011: Social book search track
    Focused Retrieval of Content and Structure, 2011
    Co-Authors: Toine Bogers, Kirstine Wilfred Christensen, Birger Larsen
    Abstract:

    In this paper, we describe our participation in the INEX 2011 Social Book Search track. We investigate the contribution of different types of Document Metadata, both social and controlled, and examine the effectiveness of re-ranking retrieval results using social features. We find that the best results are obtained using all available Document fields and topic representations.