The Experts below are selected from a list of 1992 Experts worldwide ranked by ideXlab platform

Helmut Berger - One of the best experts on this subject based on the ideXlab platform.

  • 1st international workshop on advances in patent information retrieval aspire 10
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2010
    Co-Authors: Allan Hanbury, Veronika Zenz, Helmut Berger
    Abstract:

    Patent Retrieval specialists in the 21st century face many challenges. They must search very large numbers of documents in multiple languages, expressing complex technological concepts through sophisticated legal clauses. Despite a great deal of theoretical development in Information Retrieval techniques and machine translation approaches, advanced search tools for patent professionals are still in their infancy. Patent Information Retrieval is a cross cutting research area as it contains domains such as multilingual information retrieval; language processing; image processing and retrieval; and Text Categorisation, clustering and mining. The main goal of the workshop was to gather scientists from these areas together to foster interdisciplinary collaboration and spark discussions on open topics related to search in the Intellectual Property domain. A few months before the paper submission deadline for the workshop, the IRF, supported by Matrixware, made available a collection of 400,000 patent documents in XML format for download – the AsPIRe’10 dataset. Groups submitting papers were encouraged to use this dataset for the experiments presented in the workshop papers. The AsPIRE’10 dataset was conceived to fill the gap in size between the 19 million patents in MAREC, the MAtrixware REsearch Collection, and the 20,000 patents in the One Week of MAREC collection. The AsPIRe’10 dataset continues to be available for download. Six papers were submitted to the workshop, of which five were accepted for publication after thorough review by the members of the programme committee. Three of these papers make use of the AsPIRe’10 dataset.

Hongping Hu - One of the best experts on this subject based on the ideXlab platform.

  • term frequency function of document frequency a new term weighting scheme for enterprise information retrieval
    Enterprise Information Systems, 2012
    Co-Authors: Hui Zhang, Deqing Wang, Wenjun Wu, Hongping Hu
    Abstract:

    In today's business environment, enterprises are increasingly under pressure to process the vast amount of data produced everyday within enterprises. One method is to focus on the business intelligence BI applications and increasing the commercial added-value through such business analytics activities. Term weighting scheme, which has been used to convert the documents as vectors in the term space, is a vital task in enterprise Information Retrieval IR, Text Categorisation, Text analytics, etc. When determining term weight in a document, the traditional TF-IDF scheme sets weight value for the term considering only its occurrence frequency within the document and in the entire set of documents, which leads to some meaningful terms that cannot get the appropriate weight. In this article, we propose a new term weighting scheme called Term Frequency – Function of Document Frequency TF-FDF to address this issue. Instead of using monotonically decreasing function such as Inverse Document Frequency, FDF presents a convex function that dynamically adjusts weights according to the significance of the words in a document set. This function can be manually tuned based on the distribution of the most meaningful words which semantically represent the document set. Our experiments show that the TF-FDF can achieve higher value of Normalised Discounted Cumulative Gain in IR than that of TF-IDF and its variants, and improving the accuracy of relevance ranking of the IR results.

Allan Hanbury - One of the best experts on this subject based on the ideXlab platform.

  • 1st international workshop on advances in patent information retrieval aspire 10
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2010
    Co-Authors: Allan Hanbury, Veronika Zenz, Helmut Berger
    Abstract:

    Patent Retrieval specialists in the 21st century face many challenges. They must search very large numbers of documents in multiple languages, expressing complex technological concepts through sophisticated legal clauses. Despite a great deal of theoretical development in Information Retrieval techniques and machine translation approaches, advanced search tools for patent professionals are still in their infancy. Patent Information Retrieval is a cross cutting research area as it contains domains such as multilingual information retrieval; language processing; image processing and retrieval; and Text Categorisation, clustering and mining. The main goal of the workshop was to gather scientists from these areas together to foster interdisciplinary collaboration and spark discussions on open topics related to search in the Intellectual Property domain. A few months before the paper submission deadline for the workshop, the IRF, supported by Matrixware, made available a collection of 400,000 patent documents in XML format for download – the AsPIRe’10 dataset. Groups submitting papers were encouraged to use this dataset for the experiments presented in the workshop papers. The AsPIRE’10 dataset was conceived to fill the gap in size between the 19 million patents in MAREC, the MAtrixware REsearch Collection, and the 20,000 patents in the One Week of MAREC collection. The AsPIRe’10 dataset continues to be available for download. Six papers were submitted to the workshop, of which five were accepted for publication after thorough review by the members of the programme committee. Three of these papers make use of the AsPIRe’10 dataset.

Hui Zhang - One of the best experts on this subject based on the ideXlab platform.

  • term frequency function of document frequency a new term weighting scheme for enterprise information retrieval
    Enterprise Information Systems, 2012
    Co-Authors: Hui Zhang, Deqing Wang, Wenjun Wu, Hongping Hu
    Abstract:

    In today's business environment, enterprises are increasingly under pressure to process the vast amount of data produced everyday within enterprises. One method is to focus on the business intelligence BI applications and increasing the commercial added-value through such business analytics activities. Term weighting scheme, which has been used to convert the documents as vectors in the term space, is a vital task in enterprise Information Retrieval IR, Text Categorisation, Text analytics, etc. When determining term weight in a document, the traditional TF-IDF scheme sets weight value for the term considering only its occurrence frequency within the document and in the entire set of documents, which leads to some meaningful terms that cannot get the appropriate weight. In this article, we propose a new term weighting scheme called Term Frequency – Function of Document Frequency TF-FDF to address this issue. Instead of using monotonically decreasing function such as Inverse Document Frequency, FDF presents a convex function that dynamically adjusts weights according to the significance of the words in a document set. This function can be manually tuned based on the distribution of the most meaningful words which semantically represent the document set. Our experiments show that the TF-FDF can achieve higher value of Normalised Discounted Cumulative Gain in IR than that of TF-IDF and its variants, and improving the accuracy of relevance ranking of the IR results.

John Shawetaylor - One of the best experts on this subject based on the ideXlab platform.

  • an introduction to support vector machines and other kernel based learning methods
    2000
    Co-Authors: Nello Cristianini, John Shawetaylor
    Abstract:

    From the publisher: This is the first comprehensive introduction to Support Vector Machines (SVMs), a new generation learning system based on recent advances in statistical learning theory. SVMs deliver state-of-the-art performance in real-world applications such as Text Categorisation, hand-written character recognition, image classification, biosequences analysis, etc., and are now established as one of the standard tools for machine learning and data mining. Students will find the book both stimulating and accessible, while practitioners will be guided smoothly through the material required for a good grasp of the theory and its applications. The concepts are introduced gradually in accessible and self-contained stages, while the presentation is rigorous and thorough. Pointers to relevant literature and web sites containing software ensure that it forms an ideal starting point for further study. Equally, the book and its associated web site will guide practitioners to updated literature, new applications, and on-line software.

  • an introduction to support vector machines
    Cambridge University Press (2000), 2000
    Co-Authors: Nello Cristianini, John Shawetaylor
    Abstract:

    This book is the first comprehensive introduction to Support Vector Machines (SVMs), a new generation learning system based on recent advances in statistical learning theory. The book also introduces Bayesian analysis of learning and relates SVMs to Gaussian Processes and other kernel based learning methods. SVMs deliver state-of-the-art performance in real-world applications such as Text Categorisation, hand-written character recognition, image classification, biosequences analysis, etc. Their first introduction in the early 1990s lead to a recent explosion of applications and deepening theoretical analysis, that has now established Support Vector Machines along with neural networks as one of the standard tools for machine learning and data mining. Students will find the book both stimulating and accessible, while practitioners will be guided smoothly through the material required for a good grasp of the theory and application of these techniques. The concepts are introduced gradually in accessible and self-contained stages, though in each stage the presentation is rigorous and thorough. Pointers to relevant literature and web sites containing software ensure that it forms an ideal starting point for further study. Equally the book will equip the practitioner to apply the techniques and an associated web site will provide pointers to updated literature, new applications, and on-line software.