The Experts below are selected from a list of 273 Experts worldwide ranked by ideXlab platform

Ali Selamat - One of the best experts on this subject based on the ideXlab platform.

  • parallel web crawler architecture for Clickstream Analysis
    3rd Knowledge Technology Week KTW 2011, 2012
    Co-Authors: Fatemeh Ahmadiabkenari, Ali Selamat
    Abstract:

    The tremendous growth of the Web causes many challenges for single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues. As a result, more robust algorithms needed to produce more precise and relevant search results in an appropriate timely manner. The existed Web crawlers mostly implement link dependent Web page importance metrics. One of the barriers of applying this metrics is that these metrics produce considerable communication overhead on the multi agent crawlers. Moreover, they suffer from the shortcoming of high dependency to their own index size that ends in their failure to rank Web pages with complete accuracy. Hence more enhanced metrics need to be addressed in this area. Proposing new Web page importance metric needs define a new architecture as a framework to implement the metric. The aim of this paper is to propose architecture for a focused parallel crawler. In this framework, the decision-making on Web page importance is based on a combined metric of Clickstream Analysis and context similarity Analysis to the issued queries.

  • an architecture for a focused trend parallel web crawler with the application of Clickstream Analysis
    Information Sciences, 2012
    Co-Authors: Fatemeh Ahmadiabkenari, Ali Selamat
    Abstract:

    The tremendous growth of the Web poses many challenges for all-purpose single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues regarding the enormous dimension of the World Wide Web. Hence, more enhanced and convincing algorithms are on demand to yield more precise and relevant search results in an appropriate amount of time. Since employing link based Web page importance metrics within a multi-processes crawler bears a considerable communication overhead on the overall system and cannot produce the precise answer set, employing these metrics in search engines is not an absolute solution to identify the best search answer set by the overall search system. Thus considering the employment of a link independent Web page importance metric is required to govern the priority rule within the queue of fetched URLs. The aim of this paper is to propose a modest weighted architecture for a focused structured parallel Web crawler which employs a link independent Clickstream based Web page importance metric. The experiments of this metric over the restricted boundary Web zone of our crowded UTM University Web site shows the efficiency of the proposed metric.

  • KTW - Parallel Web Crawler Architecture for Clickstream Analysis
    Communications in Computer and Information Science, 2012
    Co-Authors: Fatemeh Ahmadi-abkenari, Ali Selamat
    Abstract:

    The tremendous growth of the Web causes many challenges for single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues. As a result, more robust algorithms needed to produce more precise and relevant search results in an appropriate timely manner. The existed Web crawlers mostly implement link dependent Web page importance metrics. One of the barriers of applying this metrics is that these metrics produce considerable communication overhead on the multi agent crawlers. Moreover, they suffer from the shortcoming of high dependency to their own index size that ends in their failure to rank Web pages with complete accuracy. Hence more enhanced metrics need to be addressed in this area. Proposing new Web page importance metric needs define a new architecture as a framework to implement the metric. The aim of this paper is to propose architecture for a focused parallel crawler. In this framework, the decision-making on Web page importance is based on a combined metric of Clickstream Analysis and context similarity Analysis to the issued queries.

  • architecture for a parallel focused crawler for Clickstream Analysis
    Asian Conference on Intelligent Information and Database Systems, 2011
    Co-Authors: Ali Selamat, Fatemeh Ahmadiabkenari
    Abstract:

    The tremendous growth of the Web poses many challenges for allpurpose single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues regarding the enormous dimension of the World Wide Web. Meanwhile, more enhanced and convincing algorithms are on demand to yield more precise and relevant search results in an appropriate amount of time. Due to the fact that employing the link based Web page importance metrics in search engines is not an absolute solution to identify the best answer set by the overall search system and because employing such metrics within a multi-processes crawler bears a considerable communication overhead on the overall system, employing a link independent Web page importance metric is required to govern the priority rule within the queue of fetched URLs. The aim of this paper is to propose a modest weighted architecture for a focused structured parallel crawler in which the credit assignment to the discovered URLs is performed upon a combined metric based on Clickstream Analysis and Web page text similarity Analysis to the specified mapped topic(s).

  • ACIIDS (1) - Architecture for a parallel focused crawler for Clickstream Analysis
    Intelligent Information and Database Systems, 2011
    Co-Authors: Ali Selamat, Fatemeh Ahmadi-abkenari
    Abstract:

    The tremendous growth of the Web poses many challenges for allpurpose single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues regarding the enormous dimension of the World Wide Web. Meanwhile, more enhanced and convincing algorithms are on demand to yield more precise and relevant search results in an appropriate amount of time. Due to the fact that employing the link based Web page importance metrics in search engines is not an absolute solution to identify the best answer set by the overall search system and because employing such metrics within a multi-processes crawler bears a considerable communication overhead on the overall system, employing a link independent Web page importance metric is required to govern the priority rule within the queue of fetched URLs. The aim of this paper is to propose a modest weighted architecture for a focused structured parallel crawler in which the credit assignment to the discovered URLs is performed upon a combined metric based on Clickstream Analysis and Web page text similarity Analysis to the specified mapped topic(s).

Marcelo Perazolo - One of the best experts on this subject based on the ideXlab platform.

Fatemeh Ahmadiabkenari - One of the best experts on this subject based on the ideXlab platform.

  • parallel web crawler architecture for Clickstream Analysis
    3rd Knowledge Technology Week KTW 2011, 2012
    Co-Authors: Fatemeh Ahmadiabkenari, Ali Selamat
    Abstract:

    The tremendous growth of the Web causes many challenges for single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues. As a result, more robust algorithms needed to produce more precise and relevant search results in an appropriate timely manner. The existed Web crawlers mostly implement link dependent Web page importance metrics. One of the barriers of applying this metrics is that these metrics produce considerable communication overhead on the multi agent crawlers. Moreover, they suffer from the shortcoming of high dependency to their own index size that ends in their failure to rank Web pages with complete accuracy. Hence more enhanced metrics need to be addressed in this area. Proposing new Web page importance metric needs define a new architecture as a framework to implement the metric. The aim of this paper is to propose architecture for a focused parallel crawler. In this framework, the decision-making on Web page importance is based on a combined metric of Clickstream Analysis and context similarity Analysis to the issued queries.

  • an architecture for a focused trend parallel web crawler with the application of Clickstream Analysis
    Information Sciences, 2012
    Co-Authors: Fatemeh Ahmadiabkenari, Ali Selamat
    Abstract:

    The tremendous growth of the Web poses many challenges for all-purpose single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues regarding the enormous dimension of the World Wide Web. Hence, more enhanced and convincing algorithms are on demand to yield more precise and relevant search results in an appropriate amount of time. Since employing link based Web page importance metrics within a multi-processes crawler bears a considerable communication overhead on the overall system and cannot produce the precise answer set, employing these metrics in search engines is not an absolute solution to identify the best search answer set by the overall search system. Thus considering the employment of a link independent Web page importance metric is required to govern the priority rule within the queue of fetched URLs. The aim of this paper is to propose a modest weighted architecture for a focused structured parallel Web crawler which employs a link independent Clickstream based Web page importance metric. The experiments of this metric over the restricted boundary Web zone of our crowded UTM University Web site shows the efficiency of the proposed metric.

  • architecture for a parallel focused crawler for Clickstream Analysis
    Asian Conference on Intelligent Information and Database Systems, 2011
    Co-Authors: Ali Selamat, Fatemeh Ahmadiabkenari
    Abstract:

    The tremendous growth of the Web poses many challenges for allpurpose single-process crawlers including the presence of some irrelevant answers among search results and the coverage and scaling issues regarding the enormous dimension of the World Wide Web. Meanwhile, more enhanced and convincing algorithms are on demand to yield more precise and relevant search results in an appropriate amount of time. Due to the fact that employing the link based Web page importance metrics in search engines is not an absolute solution to identify the best answer set by the overall search system and because employing such metrics within a multi-processes crawler bears a considerable communication overhead on the overall system, employing a link independent Web page importance metric is required to govern the priority rule within the queue of fetched URLs. The aim of this paper is to propose a modest weighted architecture for a focused structured parallel crawler in which the credit assignment to the discovered URLs is performed upon a combined metric based on Clickstream Analysis and Web page text similarity Analysis to the specified mapped topic(s).

Wallace A. Pinheiro - One of the best experts on this subject based on the ideXlab platform.

Ben Y Zhao - One of the best experts on this subject based on the ideXlab platform.

  • you are how you click Clickstream Analysis for sybil detection
    USENIX Security Symposium, 2013
    Co-Authors: Gang Wang, Tristan Konolige, Christo Wilson, Xiao Wang, Haitao Zheng, Ben Y Zhao
    Abstract:

    Fake identities and Sybil accounts are pervasive in today's online communities. They are responsible for a growing number of threats, including fake product reviews, malware and spam on social networks, and astroturf political campaigns. Unfortunately, studies show that existing tools such as CAPTCHAs and graph-based Sybil detectors have not proven to be effective defenses. In this paper, we describe our work on building a practical system for detecting fake identities using server-side Clickstream models. We develop a detection approach that groups "similar" user Clickstreams into behavioral clusters, by partitioning a similarity graph that captures distances between Clickstream sequences. We validate our Clickstream models using ground-truth traces of 16,000 real and Sybil users from Renren, a large Chinese social network with 220M users. We propose a practical detection system based on these models, and show that it provides very high detection accuracy on our Clickstream traces. Finally, we worked with collaborators at Renren and LinkedIn to test our prototype on their server-side data. Following positive results, both companies have expressed strong interest in further experimentation and possible internal deployment.

  • USENIX Security Symposium - You are how you click: Clickstream Analysis for Sybil detection
    2013
    Co-Authors: Gang Wang, Tristan Konolige, Christo Wilson, Xiao Wang, Haitao Zheng, Ben Y Zhao
    Abstract:

    Fake identities and Sybil accounts are pervasive in today's online communities. They are responsible for a growing number of threats, including fake product reviews, malware and spam on social networks, and astroturf political campaigns. Unfortunately, studies show that existing tools such as CAPTCHAs and graph-based Sybil detectors have not proven to be effective defenses. In this paper, we describe our work on building a practical system for detecting fake identities using server-side Clickstream models. We develop a detection approach that groups "similar" user Clickstreams into behavioral clusters, by partitioning a similarity graph that captures distances between Clickstream sequences. We validate our Clickstream models using ground-truth traces of 16,000 real and Sybil users from Renren, a large Chinese social network with 220M users. We propose a practical detection system based on these models, and show that it provides very high detection accuracy on our Clickstream traces. Finally, we worked with collaborators at Renren and LinkedIn to test our prototype on their server-side data. Following positive results, both companies have expressed strong interest in further experimentation and possible internal deployment.