The Experts below are selected from a list of 9750 Experts worldwide ranked by ideXlab platform

James A Humenik - One of the best experts on this subject based on the ideXlab platform.

  • power source roadmaps using bibliometrics and Database tomography
    Energy, 2005
    Co-Authors: Ronald N Kostoff, Rene Tshiteya, Kirstin M Pfeil, James A Humenik, George Karypis
    Abstract:

    Database Tomography (DT) is a Textual Database analysis system consisting of two major components: (1) algorithms for extracting multi-word phrase frequencies and phrase proximities (physical closeness of the multi-word technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a Power Sources Database derived from the Science Citation Index. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the Power Sources Database, and the phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the Power Sources literature supplemented the DT results with author/journal/institution/country publication and citation data.

  • text mining using Database tomography and bibliometrics a review
    Technological Forecasting and Social Change, 2001
    Co-Authors: Ronald N Kostoff, Darrell Ray Toothman, Henry J Eberhart, James A Humenik
    Abstract:

    Abstract Database tomography (DT) is a Textual Database analysis system consisting of two major components: (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT has been used to derive technical intelligence from a variety of Textual Database sources, most recently the published technical literature as exemplified by the Science Citation Index (SCI) and the Engineering Compendex (EC). Phrase frequency analysis (the occurrence frequency of multiword technical phrases) provides the pervasive technical themes of the topical Databases of interest, and phrase proximity analysis provides the relationships among the pervasive technical themes. In the structured published literature Databases, bibliometric analysis of the Database records supplements the DT results by identifying: the recent most prolific topical area authors; the journals that contain numerous topical area papers; the institutions that produce numerous topical area papers; the keywords specified most frequently by the topical area authors; the authors whose works are cited most frequently in the topical area papers; and the particular papers and journals cited most frequently in the topical area papers. This review paper summarizes: (1) the theory and background development of DT; (2) past published and unpublished literature study results; (3) present application activities; (4) potential expansion to new DT applications. In addition, application of DT to technology forecasting is addressed.

  • fullerene data mining using bibliometrics and Database tomography
    Journal of Chemical Information and Computer Sciences, 2000
    Co-Authors: Ronald N Kostoff, Darrell Ray Toothman, Tibor Braun, Andras Schubert, James A Humenik
    Abstract:

    Database tomography (DT) is a Textual Database analysis system consisting of two major components:  (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a fullerenes Database derived from the Science Citation Index and the Engineering Compendex. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the fullerenes Database, and phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the fullerenes literature supplemented the DT results with author/journal/institution publication and citation data. Comparisons of fullerenes results with past analyses of similarly structured near-earth space, chemistry, hypersonic/supersonic flow, aircraft, and ship hydrodynam...

Bruce W Croft - One of the best experts on this subject based on the ideXlab platform.

  • sorting out searching a user interface framework for text searches
    Communications of The ACM, 1998
    Co-Authors: Ben Shneiderman, Donald Byrd, Bruce W Croft
    Abstract:

    Current user interfaces for Textual Database searching leave much to be desired: individually, they are often confusing, and as a group, they are seriously inconsistent. We propose a four-phase framework for user-interface design. The framework provides common structure and terminology for searching while preserving the distinct features of individual collections and search mechanisms.

Ronald N Kostoff - One of the best experts on this subject based on the ideXlab platform.

  • power source roadmaps using bibliometrics and Database tomography
    Energy, 2005
    Co-Authors: Ronald N Kostoff, Rene Tshiteya, Kirstin M Pfeil, James A Humenik, George Karypis
    Abstract:

    Database Tomography (DT) is a Textual Database analysis system consisting of two major components: (1) algorithms for extracting multi-word phrase frequencies and phrase proximities (physical closeness of the multi-word technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a Power Sources Database derived from the Science Citation Index. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the Power Sources Database, and the phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the Power Sources literature supplemented the DT results with author/journal/institution/country publication and citation data.

  • text mining using Database tomography and bibliometrics a review
    Technological Forecasting and Social Change, 2001
    Co-Authors: Ronald N Kostoff, Darrell Ray Toothman, Henry J Eberhart, James A Humenik
    Abstract:

    Abstract Database tomography (DT) is a Textual Database analysis system consisting of two major components: (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT has been used to derive technical intelligence from a variety of Textual Database sources, most recently the published technical literature as exemplified by the Science Citation Index (SCI) and the Engineering Compendex (EC). Phrase frequency analysis (the occurrence frequency of multiword technical phrases) provides the pervasive technical themes of the topical Databases of interest, and phrase proximity analysis provides the relationships among the pervasive technical themes. In the structured published literature Databases, bibliometric analysis of the Database records supplements the DT results by identifying: the recent most prolific topical area authors; the journals that contain numerous topical area papers; the institutions that produce numerous topical area papers; the keywords specified most frequently by the topical area authors; the authors whose works are cited most frequently in the topical area papers; and the particular papers and journals cited most frequently in the topical area papers. This review paper summarizes: (1) the theory and background development of DT; (2) past published and unpublished literature study results; (3) present application activities; (4) potential expansion to new DT applications. In addition, application of DT to technology forecasting is addressed.

  • fullerene data mining using bibliometrics and Database tomography
    Journal of Chemical Information and Computer Sciences, 2000
    Co-Authors: Ronald N Kostoff, Darrell Ray Toothman, Tibor Braun, Andras Schubert, James A Humenik
    Abstract:

    Database tomography (DT) is a Textual Database analysis system consisting of two major components:  (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large Textual Database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a fullerenes Database derived from the Science Citation Index and the Engineering Compendex. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the fullerenes Database, and phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the fullerenes literature supplemented the DT results with author/journal/institution publication and citation data. Comparisons of fullerenes results with past analyses of similarly structured near-earth space, chemistry, hypersonic/supersonic flow, aircraft, and ship hydrodynam...

Ben Shneiderman - One of the best experts on this subject based on the ideXlab platform.

  • sorting out searching a user interface framework for text searches
    Communications of The ACM, 1998
    Co-Authors: Ben Shneiderman, Donald Byrd, Bruce W Croft
    Abstract:

    Current user interfaces for Textual Database searching leave much to be desired: individually, they are often confusing, and as a group, they are seriously inconsistent. We propose a four-phase framework for user-interface design. The framework provides common structure and terminology for searching while preserving the distinct features of individual collections and search mechanisms.

  • clarifying search a user interface framework for text searches
    D-lib Magazine, 1997
    Co-Authors: Ben Shneiderman, Donald Byrd, W B Croft
    Abstract:

    Current user interfaces for Textual Database searching leave much to be desired: individually, they are often confusing, and as a group, they are seriously inconsistent. We propose a four- phase framework for user-interface design: the framework provides common structure and terminology for searching while preserving the distinct features of individual collections and search mechanisms. Users will benefit from faster learning, increased comprehension, and better control, leading to more effective searches and higher satisfaction.

Osmar R. Zaïane - One of the best experts on this subject based on the ideXlab platform.

  • Text document categorization by term association
    Icdm, 2002
    Co-Authors: M.-l. Antonie, Osmar R. Zaïane
    Abstract:

    A good text classifier is a classifier that efficiently categorizes large sets of text documents in a reasonable time frame and with an acceptable accuracy, and that provides classification rules that are human readable for possible fine-tuning. If the training of the classifier is also quick, this could become in some application domains a good asset for the classifier. Many techniques and algorithms for automatic text categorization have been devised. According to published literature, some are more accurate than others, and some provide more interpretable classification models than others. However, none can combine all the beneficial properties enumerated above. In this paper we present a novel approach for automatic text categorization that borrows from market basket analysis techniques using association rule mining in the data-mining field. We focus on two major problems: (1) finding the best term association rules in a Textual Database by generating and pruning; and (2) using the rules to build a text classifier. Our text categorization method proves to be efficient and effective, and experiments on well-known collections show that the classifier performs well. In addition, training as well as classification are both fast and the generated rules are human readable.