The Experts below are selected from a list of 6309 Experts worldwide ranked by ideXlab platform

Hai Jin - One of the best experts on this subject based on the ideXlab platform.

  • a novel Metadata Extraction method for surveillance video
    2016 International Conference on Computing Networking and Communications (ICNC), 2016
    Co-Authors: Ran Zheng, Hai Jin, Long Chen, Lei Zhu, Qin Zhang
    Abstract:

    Surveillance videos are increasing massively with monitoring cameras widely deployed in cities. Metadata of moving objects in videos can reduce video storage and be utilized in many video analysis applications. The whole Metadata Extraction is too time-consuming to be fulfilled quickly. Therefore, it is urgent to parallelize and accelerate the Extraction process. However, iterative execution of traditional Metadata Extraction methods makes the parallelization quite challenging. In this paper, we propose a novel Parallel Metadata Extraction (PME) method for surveillance video. A novel video segmentation algorithm is designed to segment whole video into independent video segments. The Metadata of each video segment is extracted independently and simultaneously on computer nodes, which will be integrated later to guarantee the completeness of the Metadata. Performance evaluations demonstrate that PME can accelerate Extraction process greatly almost without Metadata loss.

  • ICNC - A novel Metadata Extraction method for surveillance video
    2016 International Conference on Computing Networking and Communications (ICNC), 2016
    Co-Authors: Ran Zheng, Hai Jin, Long Chen, Lei Zhu, Qin Zhang
    Abstract:

    Surveillance videos are increasing massively with monitoring cameras widely deployed in cities. Metadata of moving objects in videos can reduce video storage and be utilized in many video analysis applications. The whole Metadata Extraction is too time-consuming to be fulfilled quickly. Therefore, it is urgent to parallelize and accelerate the Extraction process. However, iterative execution of traditional Metadata Extraction methods makes the parallelization quite challenging. In this paper, we propose a novel Parallel Metadata Extraction (PME) method for surveillance video. A novel video segmentation algorithm is designed to segment whole video into independent video segments. The Metadata of each video segment is extracted independently and simultaneously on computer nodes, which will be integrated later to guarantee the completeness of the Metadata. Performance evaluations demonstrate that PME can accelerate Extraction process greatly almost without Metadata loss.

  • reference Metadata Extraction from scientific papers
    Parallel and Distributed Computing: Applications and Technologies, 2011
    Co-Authors: Zhixin Guo, Hai Jin
    Abstract:

    Bibliographical information of scientific papers is of great value since the Science Citation Index is introduced to measure research impact. Most scientific documents available on the web are unstructured or semi-structured, and the automatic reference Metadata Extraction process becomes an important task. This paper describes a framework for automatic reference Metadata Extraction from scientific papers. Our system can extract title, author, journal, volume, year, and page from scientific papers in PDF. We utilize a document Metadata knowledge base to guide the reference Metadata Extraction process. The experiment results show that our system achieves a high accuracy.

  • PDCAT - Reference Metadata Extraction from scientific papers
    2011 12th International Conference on Parallel and Distributed Computing Applications and Technologies, 2011
    Co-Authors: Zhixin Guo, Hai Jin
    Abstract:

    Bibliographical information of scientific papers is of great value since the Science Citation Index is introduced to measure research impact. Most scientific documents available on the web are unstructured or semi-structured, and the automatic reference Metadata Extraction process becomes an important task. This paper describes a framework for automatic reference Metadata Extraction from scientific papers. Our system can extract title, author, journal, volume, year, and page from scientific papers in PDF. We utilize a document Metadata knowledge base to guide the reference Metadata Extraction process. The experiment results show that our system achieves a high accuracy.

  • A Rule-Based Framework of Metadata Extraction from Scientific Papers
    Distributed Computing and Applications to Business, Engineering and Science (DCABES), 2011 Tenth International Symposium on, 2011
    Co-Authors: Zhixin Guo, Hai Jin
    Abstract:

    Most scientific documents on the web are unstructured or semi-structured, and the automatic document Metadata Extraction process becomes an important task. This paper describes a framework for automatic Metadata Extraction from scientific papers. Based on a spatial and visual knowledge principle, our system can extract title, authors and abstract from scientific papers. We utilize format information such as font size and position to guide the Metadata Extraction process. The experiment results show that our system achieves a high accuracy in header Metadata Extraction which can effectively assist the automatic index creation for digital libraries.

Krisda Khankasikam - One of the best experts on this subject based on the ideXlab platform.

  • Metadata Extraction Using Case-based Reasoning for Heterogeneous Thai Documents
    2011
    Co-Authors: Krisda Khankasikam
    Abstract:

    This paper reports an experience of human-assisted process to extract Metadata from Thai documents. Nowadays, a number of Thai archives are placed online for sharing increasingly because the Internet infrastructure is powerful preserving and sharing knowledge require appropriate processes. Metadata, data about data, is a very useful information technology today because it helps users to differentiate significant from non-significant documents. The manually harvesting of these Metadata elements is highly labor-intensive, costly and time-consuming then automated is a key to successful preservation. The experiment, a prototype system by using Case-based Reasoning algorithm for Metadata Extraction is introduced. Cased-based Reasoning is an approach in artificial intelligence that differs from other approaches. The Thai Metadata Extraction were performed on some Thai articles which content related to sufficient economy and Thai folk wisdom and was evaluated the approach by using the standard precision, recall and f-measure indices. The study illustrated that this approach helps knowledge workers in a domain to come together, share educational material and greatly reduce the labor work of Metadata creation process.

  • Thai Metadata Extraction by Using Case-based Reasoning
    2010
    Co-Authors: Krisda Khankasikam
    Abstract:

    This paper reports an experience of humanassisted process to extract Metadata from Thai documents. Nowadays, a number of Thai archives are placed online for sharing increasingly because the Internet infrastructure is powerful preserving and sharing knowledge require appropriate processes. Metadata, data about data, is a very useful information technology today because it helps users to differentiate significant from non-significant documents. The manually harvesting of these Metadata elements is highly laborintensive, costly and time-consuming then automated is a key to successful preservation. The experiment, a prototype system by using Case-based Reasoning algorithm for Metadata Extraction is introduced. Casedbased Reasoning is an approach in artificial intelligence that differs from other approaches. The Thai Metadata Extraction were performed on some Thai articles which content related to sufficient economy and Thai folk wisdom and was evaluated the approach by using the standard precision, recall and f-measure indices. The study illustrated that this approach helps knowledge workers in a domain to come together, share educational material and greatly reduce the labor work of Metadata creation process.

  • Research Article Thai Metadata Extraction by using case-based reasoning การสกัุ้้ี
    2010
    Co-Authors: Krisda Khankasikam
    Abstract:

    This paper reports an experience of human-assisted process to extract Metadata from Thai documents. Nowadays, a number of Thai archives are placed online for sharing increasingly because the Internet infrastructure is powerful preserving and sharing knowledge require appropriate processes. Metadata, data about data, is a very useful information technology today because it helps users to differentiate significant documents from non-significant ones. Manually harvesting these Metadata elements is highly labor-intensive, costly and time-consuming then automated is a key to successful preservation. The experiment, a prototype system by using case-based reasoning algorithm for Metadata Extraction is introduced. Cased-based reasoning is an approach in artificial intelligence that differs from other approaches. The Thai Metadata Extraction was performed on some Thai articles which content related to sufficient economy and Thai folk wisdom and was evaluated the approach by using the standard precision, recall and f-measure indices. The study illustrated that this approach helps knowledge workers in a domain come together, share educational material and greatly reduce the labor work of Metadata creation process.

  • A Combined Template-Based and Case-Based Metadata Extraction for Heterogeneous Thai Documents
    2009 International Conference on Advanced Computer Control, 2009
    Co-Authors: Krisda Khankasikam, Nopasit Chakpitak, Thana Udomsripaiboon
    Abstract:

    Nowadays, a number of universities, laboratories, government agencies and companies that placing theirs documents online and making them searchable are increasing because the Internet infrastructure for global data access is fully functional. However, a large number of organizations have documents that lack Metadata. The lack of Metadata breaks off not only the discovery and dissemination of these documents over the Internet, but also their connectivity with other documents. Unfortunately, manual Metadata Extraction is expensive and time-consuming for a large document, and most existing automated Metadata Extraction approaches have focused on specific domains and homogeneous documents. In this paper, we propose a combined cased-based and template-based Metadata Extraction approach to solve these issues. The key idea of solving the heterogeneity is to classify documents into equivalent groups so that each document group contains similar documents only. Next, for each document group we have a template of previous case that contains a process to extract Metadata from documents in the group.

  • A Unified Framework for Thai Metadata Extraction Using Case-Based Reasoning
    2008 International Conference on Advanced Computer Theory and Engineering, 2008
    Co-Authors: Krisda Khankasikam, Nopasit Chakpitak
    Abstract:

    Metadata is a very popular word in information technology today because it helps users to differentiate significant documents from non-significant documents. With the growth of the Internet and related tools, there has been a rapid growth of online resources. However, lack of Metadata available for these resources stops their discovery and dissemination over the Internet. The process for manual Metadata Extraction is time-consuming, costly, and labor-extensive. This paper describes a framework for automatic Metadata Extraction from electronic Thai documents. The system consists of three main components: a case retrieval module for comparing problem case and stored case using nearest neighbor retrieval technique, a Metadata creation module for automatically extracting Metadata from electronic Thai documents using Thai information Extraction techniques, and a Metadata verification module for correcting the errors in extracted Metadata. The experimental results show that using the proposed framework could reduce the labor work of Thai Metadata creation process.

Eren Manavoglu - One of the best experts on this subject based on the ideXlab platform.

  • rule based word clustering for document Metadata Extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Lee C Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science document search engine, uses automatic text classification methods for document indexing. Text classification uses a document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields Extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of document Metadata Extraction.

  • SAC - Rule-based word clustering for document Metadata Extraction
    Proceedings of the 2005 ACM symposium on Applied computing - SAC '05, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, C. Lee Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science document search engine, uses automatic text classification methods for document indexing. Text classification uses a document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields Extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of document Metadata Extraction.

  • automatic document Metadata Extraction using support vector machines
    ACM IEEE Joint Conference on Digital Libraries, 2003
    Co-Authors: C L Giles, Eren Manavoglu, Zhenyue Zhang
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata Extraction. We describe a support vector machine classification-based method for Metadata Extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata Extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata Extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata Extraction method was originally designed to improve the Metadata Extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

  • JCDL - Automatic document Metadata Extraction using support vector machines
    2003 Joint Conference on Digital Libraries 2003. Proceedings., 2003
    Co-Authors: Hui Han, C L Giles, Eren Manavoglu, Zhenyue Zhang, Hongyuan Zha, Edward A. Fox
    Abstract:

    Automatic Metadata generation provides scalability and usability for digital libraries and their collections. Machine learning methods offer robust and adaptable automatic Metadata Extraction. We describe a support vector machine classification-based method for Metadata Extraction from header part of research papers and show that it outperforms other machine learning methods on the same task. The method first classifies each line of the header into one or more of 15 classes. An iterative convergence procedure is then used to improve the line classification by using the predicted class labels of its neighbor lines in the previous round. Further Metadata Extraction is done by seeking the best chunk boundaries of each line. We found that discovery and use of the structural patterns of the data and domain based word clustering can improve the Metadata Extraction performance. An appropriate feature normalization also greatly improves the classification performance. Our Metadata Extraction method was originally designed to improve the Metadata Extraction quality of the digital libraries Citeseer [S. Lawrence et al., (1999)] and EbizSearch [Y. Petinot et al., (2003)]. We believe it can be generalized to other digital libraries.

Lee C Giles - One of the best experts on this subject based on the ideXlab platform.

  • crowd sourcing web knowledge for Metadata Extraction
    ACM IEEE Joint Conference on Digital Libraries, 2014
    Co-Authors: Wenyi Huang, Chen Liang, Lee C Giles
    Abstract:

    We explore a new Metadata Extraction framework without human annotators with the ground truth harvested from Web. A new training sample is selected based on not only the uncertainty and representativeness in the unlabeled pool, but also on its availability and credibility in Web knowledge bases. We construct a dataset of 4329 books with valid Metadata and evaluate our approach using 5 Web book databases as oracles. Empirical results demonstrate its effectiveness and efficiency.

  • rule based word clustering for document Metadata Extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Lee C Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science document search engine, uses automatic text classification methods for document indexing. Text classification uses a document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields Extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of document Metadata Extraction.

Xiangmin Zhang - One of the best experts on this subject based on the ideXlab platform.

  • rule based word clustering for document Metadata Extraction
    ACM Symposium on Applied Computing, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, Lee C Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science document search engine, uses automatic text classification methods for document indexing. Text classification uses a document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields Extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of document Metadata Extraction.

  • SAC - Rule-based word clustering for document Metadata Extraction
    Proceedings of the 2005 ACM symposium on Applied computing - SAC '05, 2005
    Co-Authors: Hui Han, Eren Manavoglu, Hongyuan Zha, C. Lee Giles, Kostas Tsioutsiouliklis, Xiangmin Zhang
    Abstract:

    Text classification is still an important problem for unlabeled text; CiteSeer, a computer science document search engine, uses automatic text classification methods for document indexing. Text classification uses a document's original text words as the primary feature representation. However, such representation usually comes with high dimensionality and feature sparseness. Word clustering is an effective approach to reduce feature dimensionality and feature sparseness, and improve text classification performance. This paper introduces a domain Rule-based word clustering method for cluster feature representation. The clusters are formed from various domain databases and the word orthographic properties. Besides significant dimensionality reduction, such cluster feature representations show a 6.6% absolute improvement on average on classification performance of document header lines and a 8.4% absolute improvement on the overall accuracy of bibliographic fields Extraction, in contrast to feature representation just based on the original text words. Our word clustering even outperforms the distributional word clustering in the context of document Metadata Extraction.