The Experts below are selected from a list of 303 Experts worldwide ranked by ideXlab platform

Eric Ziecker - One of the best experts on this subject based on the ideXlab platform.

Shojiro Nishio - One of the best experts on this subject based on the ideXlab platform.

  • wikipedia mining for huge scale japanese association Thesaurus Construction
    Advanced Information Networking and Applications, 2008
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Masahiro Ito, Shojiro Nishio
    Abstract:

    Wikipedia, a huge scale Web-based dictionary, is an impressive corpus for knowledge extraction. We already proved that Wikipedia can be used for constructing an English association Thesaurus and our link structure mining method is significantly effective for this aim. However, we want to find out how we can apply this method to other languages and what the requirements, differences and characteristics are. Nowadays, Wikipedia supports more than 250 languages such as English, German, French, Polish and Japanese. Among Asian languages, the Japanese Wikipedia is the largest corpus in Wikipedia. In this research, therefore, we analyzed all Japanese articles in Wikipedia and constructed a huge scale Japanese association Thesaurus. After constructing the Thesaurus, we realized that it shows several impressive characteristics depending on language and culture.

  • wikipedia mining for an association web Thesaurus Construction
    Web Information Systems Engineering, 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Wikipedia has become a huge phenomenon on the WWW. As a corpus for knowledge extraction, it has various impressive characteristics such as a huge amount of articles, live updates, a dense link structure, brief link texts and URL identification for concepts. In this paper, we propose an efficient link mining method pfibf (Path Frequency - Inversed Backward link Frequency) and the extension method "forward / backward link weighting (FB weighting)" in order to construct a huge scale association Thesaurus. We proved the effectiveness of our proposed methods compared with other conventional methods such as cooccurrence analysis and TF-IDF.

  • a Thesaurus Construction method from large scaleweb dictionaries
    Advanced Information Networking and Applications, 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Web-based dictionaries, such as Wikipedia, have become dramatically popular among the Internet users in past several years. The important characteristic of Web-based dictionary is not only the huge amount of articles, but also hyperlinks. Hyperlinks have various information more than just providing transfer function between pages. In this paper, we propose an efficient method to analyze the link structure of Web-based dictionaries to construct an association Thesaurus. We have already applied it to Wikipedia, a huge scale Web-based dictionary which has a dense link structure, as a corpus. We developed a search engine for evaluation, then conducted a number of experiments to compare our method with other traditional methods such as cooccurrence analysis.

  • AINA - A Thesaurus Construction Method from Large ScaleWeb Dictionaries
    21st International Conference on Advanced Information Networking and Applications (AINA '07), 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Web-based dictionaries, such as Wikipedia, have become dramatically popular among the Internet users in past several years. The important characteristic of Web-based dictionary is not only the huge amount of articles, but also hyperlinks. Hyperlinks have various information more than just providing transfer function between pages. In this paper, we propose an efficient method to analyze the link structure of Web-based dictionaries to construct an association Thesaurus. We have already applied it to Wikipedia, a huge scale Web-based dictionary which has a dense link structure, as a corpus. We developed a search engine for evaluation, then conducted a number of experiments to compare our method with other traditional methods such as cooccurrence analysis.

  • WISE - Wikipedia mining for an association web Thesaurus Construction
    Web Information Systems Engineering – WISE 2007, 1
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Wikipedia has become a huge phenomenon on the WWW. As a corpus for knowledge extraction, it has various impressive characteristics such as a huge amount of articles, live updates, a dense link structure, brief link texts and URL identification for concepts. In this paper, we propose an efficient link mining method pfibf (Path Frequency - Inversed Backward link Frequency) and the extension method "forward / backward link weighting (FB weighting)" in order to construct a huge scale association Thesaurus. We proved the effectiveness of our proposed methods compared with other conventional methods such as cooccurrence analysis and TF-IDF.

Louise F. Spiteri - One of the best experts on this subject based on the ideXlab platform.

  • Word Association Testing and Thesaurus Construction
    Proceedings of the Annual Conference of CAIS Actes du congrès annuel de l'ACSI, 2013
    Co-Authors: Louise F. Spiteri
    Abstract:

    This paper examines the suitability of word association tests to generate user-derived descriptors, descriptor hierarchies, and categories of inter-term relationships. The typical assumption underlying these word association tests is that the response terms function either as synonyms or antonyms, an assumption that restricts unnecessarily the potential value of such tests. Rather than assuming how people inter-relate two terms, it may be more useful to ask participants to explain why they think these two terms are related. In this study, thirty library and information science practitioners were asked to provide as many response words as they could for fifteen stimulus terms and to describe how the response and stimulus terms were inter-related. The word association test was successful in generating a set of user-derived descriptors . Participants identified twenty types of inter-term relationships, the most commonly-cited of which are type, part, synonym, activity, and tool. That the participants identified a total of twenty types of relationships suggests also that word association tests can serve as a valuable tool in examining the different ways users group terms and the types of inter-term relationships that end users most commonly associate with any given concept and its response terms.

  • Word Association Testing and Thesaurus Construction: A Pilot Study
    Cataloging & Classification Quarterly, 2005
    Co-Authors: Louise F. Spiteri
    Abstract:

    This pilot study examines the use of word association testing in the derivation of user-derived descriptors, descriptor hierarchies, and categories of inter-term relationships for the purpose of Thesaurus Construction. Ten participants, who were students, were presented with a test-bed of 15 domain-specific stimulus terms and were asked to provide as many response words as they could for each stimulus term and to describe how the response and stimulus terms are inter-related. The word association test was successful in generating a significant number of word pairs and facet indicators that could be used to display inter-term relationships in thesauri.

  • Word association testing and Thesaurus Construction: Defining inter-term relationships
    2001
    Co-Authors: Louise F. Spiteri
    Abstract:

    This paper presents a theoretical framework that incorporates word association testing into the design of hierarchical displays in information retrieval (IR) thesauri. Word association tests typically present participants with a test bed of single-concept terms ("Stimulus terms") from one subject domain; for each stimulus term, the participants are asked to write down as many words. . .

  • The use of facet analysis in Information Retrieval Thesauri: an examination of selected guidelines for Thesaurus Construction
    Cataloging & Classification Quarterly, 1998
    Co-Authors: Louise F. Spiteri
    Abstract:

    ABSTRACT Facet analysis has been used in the Construction of faceted thesauri since the publication of the Information Retrieval Thesaurus of Education Terms in 1968. In spite of the growth in the number of faceted thesauri since then, there appears to be little consensus among Thesaurus designers regarding how the principles of facet analysis are to be used in thesauri. An examination of various national and international guidelines for Thesaurus Construction reveals that they emphasize primarily the Construction of alphabetical thesauri, but provide little guidance in the use of facet analysis in thesauri.

Bokyung Yang - One of the best experts on this subject based on the ideXlab platform.

  • experiments in automatic statistical Thesaurus Construction
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 1992
    Co-Authors: Carolyn J Crouch, Bokyung Yang
    Abstract:

    A well constructed Thesaurus has long been recognized as a valuable tool in the effective operation of an information retrieval system. This paper reports the results of experiments designed to determine the validity of an approach to the automatic Construction of global thesauri (described originally by Crouch in [1] and [2] based on a clustering of the document collection. The authors validate the approach by showing that the use of thesauri generated by this method results in substantial improvements in retrieval effectiveness in four test collections. The term discrimination value theory, used in the Thesaurus generation algorithm to determine a term's membership in a particular Thesaurus class, is found not to be useful in distinguishing a “good” from an “indifferent” or “poor” Thesaurus class). In conclusion, the authors suggest an alternate approach to automatic Thesaurus Construction which greatly simplifies the work of producing viable Thesaurus classes. Experimental results show that the alternate approach described herein in some cases produces thesauri which are comparable in retrieval effectiveness to those produced by the first method at much lower cost.

  • SIGIR - Experiments in automatic statistical Thesaurus Construction
    Proceedings of the 15th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '92, 1992
    Co-Authors: Carolyn J Crouch, Bokyung Yang
    Abstract:

    A well constructed Thesaurus has long been recognized as a valuable tool in the effective operation of an information retrieval system. This paper reports the results of experiments designed to determine the validity of an approach to the automatic Construction of global thesauri (described originally by Crouch in [1] and [2] based on a clustering of the document collection. The authors validate the approach by showing that the use of thesauri generated by this method results in substantial improvements in retrieval effectiveness in four test collections. The term discrimination value theory, used in the Thesaurus generation algorithm to determine a term's membership in a particular Thesaurus class, is found not to be useful in distinguishing a “good” from an “indifferent” or “poor” Thesaurus class). In conclusion, the authors suggest an alternate approach to automatic Thesaurus Construction which greatly simplifies the work of producing viable Thesaurus classes. Experimental results show that the alternate approach described herein in some cases produces thesauri which are comparable in retrieval effectiveness to those produced by the first method at much lower cost.

Kotaro Nakayama - One of the best experts on this subject based on the ideXlab platform.

  • wikipedia mining for huge scale japanese association Thesaurus Construction
    Advanced Information Networking and Applications, 2008
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Masahiro Ito, Shojiro Nishio
    Abstract:

    Wikipedia, a huge scale Web-based dictionary, is an impressive corpus for knowledge extraction. We already proved that Wikipedia can be used for constructing an English association Thesaurus and our link structure mining method is significantly effective for this aim. However, we want to find out how we can apply this method to other languages and what the requirements, differences and characteristics are. Nowadays, Wikipedia supports more than 250 languages such as English, German, French, Polish and Japanese. Among Asian languages, the Japanese Wikipedia is the largest corpus in Wikipedia. In this research, therefore, we analyzed all Japanese articles in Wikipedia and constructed a huge scale Japanese association Thesaurus. After constructing the Thesaurus, we realized that it shows several impressive characteristics depending on language and culture.

  • wikipedia mining for an association web Thesaurus Construction
    Web Information Systems Engineering, 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Wikipedia has become a huge phenomenon on the WWW. As a corpus for knowledge extraction, it has various impressive characteristics such as a huge amount of articles, live updates, a dense link structure, brief link texts and URL identification for concepts. In this paper, we propose an efficient link mining method pfibf (Path Frequency - Inversed Backward link Frequency) and the extension method "forward / backward link weighting (FB weighting)" in order to construct a huge scale association Thesaurus. We proved the effectiveness of our proposed methods compared with other conventional methods such as cooccurrence analysis and TF-IDF.

  • a Thesaurus Construction method from large scaleweb dictionaries
    Advanced Information Networking and Applications, 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Web-based dictionaries, such as Wikipedia, have become dramatically popular among the Internet users in past several years. The important characteristic of Web-based dictionary is not only the huge amount of articles, but also hyperlinks. Hyperlinks have various information more than just providing transfer function between pages. In this paper, we propose an efficient method to analyze the link structure of Web-based dictionaries to construct an association Thesaurus. We have already applied it to Wikipedia, a huge scale Web-based dictionary which has a dense link structure, as a corpus. We developed a search engine for evaluation, then conducted a number of experiments to compare our method with other traditional methods such as cooccurrence analysis.

  • AINA - A Thesaurus Construction Method from Large ScaleWeb Dictionaries
    21st International Conference on Advanced Information Networking and Applications (AINA '07), 2007
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Web-based dictionaries, such as Wikipedia, have become dramatically popular among the Internet users in past several years. The important characteristic of Web-based dictionary is not only the huge amount of articles, but also hyperlinks. Hyperlinks have various information more than just providing transfer function between pages. In this paper, we propose an efficient method to analyze the link structure of Web-based dictionaries to construct an association Thesaurus. We have already applied it to Wikipedia, a huge scale Web-based dictionary which has a dense link structure, as a corpus. We developed a search engine for evaluation, then conducted a number of experiments to compare our method with other traditional methods such as cooccurrence analysis.

  • WISE - Wikipedia mining for an association web Thesaurus Construction
    Web Information Systems Engineering – WISE 2007, 1
    Co-Authors: Kotaro Nakayama, Takahiro Hara, Shojiro Nishio
    Abstract:

    Wikipedia has become a huge phenomenon on the WWW. As a corpus for knowledge extraction, it has various impressive characteristics such as a huge amount of articles, live updates, a dense link structure, brief link texts and URL identification for concepts. In this paper, we propose an efficient link mining method pfibf (Path Frequency - Inversed Backward link Frequency) and the extension method "forward / backward link weighting (FB weighting)" in order to construct a huge scale association Thesaurus. We proved the effectiveness of our proposed methods compared with other conventional methods such as cooccurrence analysis and TF-IDF.