The Experts below are selected from a list of 249 Experts worldwide ranked by ideXlab platform

A. Naveenkumar - One of the best experts on this subject based on the ideXlab platform.

  • A Survey on the Classification of Dark Web using Unclassified Ontology Method
    Data mining and knowledge engineering, 2012
    Co-Authors: M. Sreekrishna, B. Chitra, A. Naveenkumar
    Abstract:

    The deep web are the web that are not a part of surface web. Due to the large volume of data deep web have grained a large attention in recent years. Traditional search engines cannot be used to retrieve content in the deep Web. Those pages do not exist until they are created dynamically as the result of a specific search. The deep web is found to be large magnitude than the surface web. Further those deep web mostly comprises of online domain specific databases, which are accessed by using web query interfaces. In order to make the extraction relevant to user it is necessary to classify the deep web database. In this paper unclassified ontology based web classification method is used for to classify the data in the deep web. This method involves completely unclassified set of data and uses Wikipedia Category Network for to analyze the meta-information of the deep web sources. The result of the experiment is found to more accurate and fine-grained classification when compared to the existing approaches.

Michael Strube - One of the best experts on this subject based on the ideXlab platform.

  • decoding wikipedia categories for knowledge acquisition
    National Conference on Artificial Intelligence, 2008
    Co-Authors: Vivi Nastase, Michael Strube
    Abstract:

    This paper presents an approach to acquire knowledge from Wikipedia categories and the Category Network. Many Wikipedia categories have complex names which reflect human classification and organizing instances, and thus encode knowledge about class attributes, taxonomic and other semantic relations. We decode the names and refer back to the Network to induce relations between concepts in Wikipedia represented through pages or categories. The Category structure allows us to propagate a relation detected between constituents of a Category name to numerous concept links. The results of the process are evaluated against ResearchCyc and a subset also by human judges. The results support the idea that Wikipedia Category names are a rich source of useful and accurate knowledge.

  • AAAI - Decoding wikipedia categories for knowledge acquisition
    2008
    Co-Authors: Vivi Nastase, Michael Strube
    Abstract:

    This paper presents an approach to acquire knowledge from Wikipedia categories and the Category Network. Many Wikipedia categories have complex names which reflect human classification and organizing instances, and thus encode knowledge about class attributes, taxonomic and other semantic relations. We decode the names and refer back to the Network to induce relations between concepts in Wikipedia represented through pages or categories. The Category structure allows us to propagate a relation detected between constituents of a Category name to numerous concept links. The results of the process are evaluated against ResearchCyc and a subset also by human judges. The results support the idea that Wikipedia Category names are a rich source of useful and accurate knowledge.

  • ESWC - Distinguishing between instances and classes in the wikipedia taxonomy
    Lecture Notes in Computer Science, 1
    Co-Authors: Cäcilia Zirn, Vivi Nastase, Michael Strube
    Abstract:

    This paper presents an automatic method for differentiating between instances and classes in a large scale taxonomy induced from the Wikipedia Category Network. The method exploits characteristics of the Category names and the structure of the Network. The approach we present is the first attempt to make this distinction automatically in a large scale resource. In contrast, this distinction has been made in WordNet and Cyc based on manual annotations. The result of the process is evaluated against ResearchCyc. On the subNetwork shared by our taxonomy and ResearchCyc we report 84.52% accuracy.

Marius Pasca - One of the best experts on this subject based on the ideXlab platform.

  • WSDM - German Typographers vs. German Grammar: Decomposition of Wikipedia Category Labels into Attribute-Value Pairs
    Proceedings of the Tenth ACM International Conference on Web Search and Data Mining, 2017
    Co-Authors: Marius Pasca
    Abstract:

    Given an instance (Julieta Pinto), most methods for open-domain information extraction focus on acquiring knowledge in the form of either class labels (Costa Rican short story writers, Women novelists) referring to concepts to which the instance belongs; or facts (nationality: Costa Rica) connecting the instance (Julieta Pinto) to other instances or concepts (Costa Rica), where the fact and the other instance often take the form of an attribute (nationality) and a value (Costa Rica) respectively. From extraction through internal representation and storage, class labels and facts are treated as if they carved out disconnected slices within the larger space of factual knowledge. This paper argues that class labels and facts pertaining to an instance exist in symbiosis rather than as a dichotomy. A constituent (Costa Rican) within a class label (Costa Rican short story writers) of an instance may be indicative of a fact (nationality: Costa Rica) applicable to the instance and vice-versa. As an illustration of the relationship between class labels and facts, the paper introduces an open-domain method for the better understanding of the semantics of class labels in one of the larger and most widely-used repositories of knowledge, namely the categories in the Wikipedia Category Network. The method exploits the Category Network to associate constituents (Costa Rican) within names of Wikipedia categories, with attributes (nationality) that explain their role.

  • COLING - Revisiting Taxonomy Induction over Wikipedia
    2016
    Co-Authors: Amit Gupta, Francesco Piccinno, Mikhail Kozhevnikov, Marius Pasca, Daniele Pighin
    Abstract:

    Guided by multiple heuristics, a unified taxonomy of entities and categories is distilled from the Wikipedia Category Network. A comprehensive evaluation, based on the analysis of upward generalization paths, demonstrates that the taxonomy supports generalizations which are more than twice as accurate as the state of the art. The taxonomy is available at http://headstaxonomy.com.

Rabih Bashroush - One of the best experts on this subject based on the ideXlab platform.

  • Corporate Information Security Investment Decisions
    Research Anthology on Artificial Intelligence Applications in Security, 2021
    Co-Authors: Daniel Schatz, Rabih Bashroush
    Abstract:

    This article describes how with information security steadily moving up on board room agendas, security programs are found to be under increasing scrutiny by practitioners. This level of attention by senior business leaders is new to many security professionals as their field has been of limited interest to non-executive directors so far. Currently, they have to regularly report on efficiency and value of their security capabilities whilst being measured against business priorities. Based on the Grounded Theory approach, the authors analysed the data gathered in a series of interviews with senior professionals in order to identify key factors in the context of information security investment decisions. The authors present detailed findings in context of a simplified framework that security practitioners can utilise for critical review or improvements of investment decisions in their own environments. Extensive details for each Category as extracted through a qualitative data analysis are provided along with a Category Network analysis that highlights strong relationships within the framework.

  • Corporate Information Security Investment Decisions: A Qualitative Data Analysis Approach
    International Journal of Enterprise Information Systems, 2018
    Co-Authors: Daniel Schatz, Rabih Bashroush
    Abstract:

    This article describes how with information security steadily moving up on board room agendas, security programs are found to be under increasing scrutiny by practitioners. This level of attention by senior business leaders is new to many security professionals as their field has been of limited interest to non-executive directors so far. Currently, they have to regularly report on efficiency and value of their security capabilities whilst being measured against business priorities. Based on the Grounded Theory approach, the authors analysed the data gathered in a series of interviews with senior professionals in order to identify key factors in the context of information security investment decisions. The authors present detailed findings in context of a simplified framework that security practitioners can utilise for critical review or improvements of investment decisions in their own environments. Extensive details for each Category as extracted through a qualitative data analysis are provided along with a Category Network analysis that highlights strong relationships within the framework.

Panagiotis Papapetrou - One of the best experts on this subject based on the ideXlab platform.

  • analysis of cluster structure in large scale english wikipedia Category Networks
    Intelligent Data Analysis, 2013
    Co-Authors: Thidawan Klaysri, Oded Lachish, Trevor Fenner, Mark Levene, Panagiotis Papapetrou
    Abstract:

    In this paper we propose a framework for analysing the structure of a large-scale social media Network, a topic of significant recent interest. Our study is focused on the Wikipedia Category Network, where nodes correspond to Wikipedia categories and edges connect two nodes if the nodes share at least one common page within the Wikipedia Network. Moreover, each edge is given a weight that corresponds to the number of pages shared between the two categories that it connects. We study the structure of Category clusters within the three complete English Wikipedia Category Networks from 2010 to 2012. We observe that Category clusters appear in the form of well-connected components that are naturally clustered together. For each dataset we obtain a graph, which we call the t-filtered Category graph, by retaining just a single edge linking each pair of categories for which the weight of the edge exceeds some specified threshold t. Our framework exploits this graph structure and identifies connected components within the t-filtered Category graph. We studied the large-scale structural properties of the three Wikipedia Category Networks using the proposed approach. We found that the number of categories, the number of clusters of size two, and the size of the largest cluster within the graph all appear to follow power laws in the threshold t. Furthermore, for each Network we found the value of the threshold t for which increasing the threshold to t+1 caused the "giant" largest cluster to diffuse into two or more smaller clusters of significant size and studied the semantics behind this diffusion.

  • IDA - Analysis of Cluster Structure in Large-Scale English Wikipedia Category Networks
    Advances in Intelligent Data Analysis XII, 2013
    Co-Authors: Thidawan Klaysri, Oded Lachish, Trevor Fenner, Mark Levene, Panagiotis Papapetrou
    Abstract:

    In this paper we propose a framework for analysing the structure of a large-scale social media Network, a topic of significant recent interest. Our study is focused on the Wikipedia Category Network, where nodes correspond to Wikipedia categories and edges connect two nodes if the nodes share at least one common page within the Wikipedia Network. Moreover, each edge is given a weight that corresponds to the number of pages shared between the two categories that it connects. We study the structure of Category clusters within the three complete English Wikipedia Category Networks from 2010 to 2012. We observe that Category clusters appear in the form of well-connected components that are naturally clustered together. For each dataset we obtain a graph, which we call the t-filtered Category graph, by retaining just a single edge linking each pair of categories for which the weight of the edge exceeds some specified threshold t. Our framework exploits this graph structure and identifies connected components within the t-filtered Category graph. We studied the large-scale structural properties of the three Wikipedia Category Networks using the proposed approach. We found that the number of categories, the number of clusters of size two, and the size of the largest cluster within the graph all appear to follow power laws in the threshold t. Furthermore, for each Network we found the value of the threshold t for which increasing the threshold to t+1 caused the "giant" largest cluster to diffuse into two or more smaller clusters of significant size and studied the semantics behind this diffusion.