The Experts below are selected from a list of 5397 Experts worldwide ranked by ideXlab platform

Bela Gipp - One of the best experts on this subject based on the ideXlab platform.

  • Academic Plagiarism Detection: A Systematic Literature Review
    ACM Computing Surveys, 2020
    Co-Authors: Tomáš Foltýnek, Norman Meuschke, Bela Gipp
    Abstract:

    This article summarizes the research on computational methods to detect academic Plagiarism by systematically reviewing 239 research papers published between 2013 and 2018. To structure the presentation of the research contributions, we propose novel technically oriented typologies for Plagiarism prevention and Detection efforts, the forms of academic Plagiarism, and computational Plagiarism Detection methods. We show that academic Plagiarism Detection is a highly active research field. Over the period we review, the field has seen major advances regarding the automated Detection of strongly obfuscated and thus hard-to-identify forms of academic Plagiarism. These improvements mainly originate from better semantic text analysis methods, the investigation of non-textual content features, and the application of machine learning. We identify a research gap in the lack of methodologically thorough performance evaluations of Plagiarism Detection systems. Concluding from our analysis, we see the integration of heterogeneous analysis methods for textual and non-textual content features using machine learning as the most promising area for future research contributions to improve the Detection of academic Plagiarism further.

  • JCDL - A First Step Towards Content Protecting Plagiarism Detection
    Proceedings of the ACM IEEE Joint Conference on Digital Libraries in 2020, 2020
    Co-Authors: Cornelius Ihle, Moritz Schubotz, Norman Meuschke, Bela Gipp
    Abstract:

    Plagiarism Detection systems are essential tools for safeguarding academic and educational integrity. However, today's systems require disclosing the full content of the input documents and the document collection to which the input documents are compared. Moreover, the systems are centralized and under the control of individual, typically commercial providers. This situation raises procedural and legal concerns regarding the confidentiality of sensitive data, which can limit or prohibit the use of Plagiarism Detection services. To eliminate these weaknesses of current systems, we seek to devise a Plagiarism Detection approach that does not require a centralized provider nor exposing any content as cleartext. This paper presents the initial results of our research. Specifically, we employ Private Set Intersection to devise a content-protecting variant of the citation-based similarity measure Bibliographic Coupling implemented in our Plagiarism Detection system HyPlag. Our evaluation shows that the content-protecting method achieves the same Detection effectiveness as the original method while making common attacks to disclose the protected content practically infeasible. Our future work will extend this successful proof-of-concept by devising Plagiarism Detection methods that can analyze the entire content of documents without disclosing it as cleartext.

  • Citation-based Plagiarism Detection - Citation-based Plagiarism Detection
    Citation-based Plagiarism Detection, 2014
    Co-Authors: Bela Gipp
    Abstract:

    When the author first considered the use of citation information as a method to detect Plagiarism, he assumed this concept had already been explored or even integrated into today’s Plagiarism Detection systems (PDS). After all, citations and references of scholarly publications have long been recognized as containing valuable semantic relatedness information for documents, as demonstrated in Section 3.2.

  • citation based Plagiarism Detection
    2014
    Co-Authors: Bela Gipp
    Abstract:

    When the author first considered the use of citation information as a method to detect Plagiarism, he assumed this concept had already been explored or even integrated into today’s Plagiarism Detection systems (PDS). After all, citations and references of scholarly publications have long been recognized as containing valuable semantic relatedness information for documents, as demonstrated in Section 3.2.

  • demonstration of citation pattern analysis for Plagiarism Detection
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2013
    Co-Authors: Bela Gipp, Norman Meuschke, Corinna Breitinger, Mario Lipinski, Andreas Nurnberger
    Abstract:

    Limitations of Plagiarism Detection Systems State-of-the-art Plagiarism Detection approaches capably identify copy & paste and to some extent slightly modified Plagiarism. However, they cannot reliably identify strongly disguised Plagiarism forms, including paraphrases, translated Plagiarism, and idea Plagiarism, which are forms of Plagiarism more commonly found in scientific texts. This weakness of current systems results in a large fraction of today’s scientific Plagiarism going undetected.

Norman Meuschke - One of the best experts on this subject based on the ideXlab platform.

  • Academic Plagiarism Detection: A Systematic Literature Review
    ACM Computing Surveys, 2020
    Co-Authors: Tomáš Foltýnek, Norman Meuschke, Bela Gipp
    Abstract:

    This article summarizes the research on computational methods to detect academic Plagiarism by systematically reviewing 239 research papers published between 2013 and 2018. To structure the presentation of the research contributions, we propose novel technically oriented typologies for Plagiarism prevention and Detection efforts, the forms of academic Plagiarism, and computational Plagiarism Detection methods. We show that academic Plagiarism Detection is a highly active research field. Over the period we review, the field has seen major advances regarding the automated Detection of strongly obfuscated and thus hard-to-identify forms of academic Plagiarism. These improvements mainly originate from better semantic text analysis methods, the investigation of non-textual content features, and the application of machine learning. We identify a research gap in the lack of methodologically thorough performance evaluations of Plagiarism Detection systems. Concluding from our analysis, we see the integration of heterogeneous analysis methods for textual and non-textual content features using machine learning as the most promising area for future research contributions to improve the Detection of academic Plagiarism further.

  • JCDL - A First Step Towards Content Protecting Plagiarism Detection
    Proceedings of the ACM IEEE Joint Conference on Digital Libraries in 2020, 2020
    Co-Authors: Cornelius Ihle, Moritz Schubotz, Norman Meuschke, Bela Gipp
    Abstract:

    Plagiarism Detection systems are essential tools for safeguarding academic and educational integrity. However, today's systems require disclosing the full content of the input documents and the document collection to which the input documents are compared. Moreover, the systems are centralized and under the control of individual, typically commercial providers. This situation raises procedural and legal concerns regarding the confidentiality of sensitive data, which can limit or prohibit the use of Plagiarism Detection services. To eliminate these weaknesses of current systems, we seek to devise a Plagiarism Detection approach that does not require a centralized provider nor exposing any content as cleartext. This paper presents the initial results of our research. Specifically, we employ Private Set Intersection to devise a content-protecting variant of the citation-based similarity measure Bibliographic Coupling implemented in our Plagiarism Detection system HyPlag. Our evaluation shows that the content-protecting method achieves the same Detection effectiveness as the original method while making common attacks to disclose the protected content practically infeasible. Our future work will extend this successful proof-of-concept by devising Plagiarism Detection methods that can analyze the entire content of documents without disclosing it as cleartext.

  • demonstration of citation pattern analysis for Plagiarism Detection
    International ACM SIGIR Conference on Research and Development in Information Retrieval, 2013
    Co-Authors: Bela Gipp, Norman Meuschke, Corinna Breitinger, Mario Lipinski, Andreas Nurnberger
    Abstract:

    Limitations of Plagiarism Detection Systems State-of-the-art Plagiarism Detection approaches capably identify copy & paste and to some extent slightly modified Plagiarism. However, they cannot reliably identify strongly disguised Plagiarism forms, including paraphrases, translated Plagiarism, and idea Plagiarism, which are forms of Plagiarism more commonly found in scientific texts. This weakness of current systems results in a large fraction of today’s scientific Plagiarism going undetected.

  • SIGIR - Demonstration of citation pattern analysis for Plagiarism Detection
    Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval - SIGIR '13, 2013
    Co-Authors: Bela Gipp, Norman Meuschke, Corinna Breitinger, Mario Lipinski, Andreas Nurnberger
    Abstract:

    Limitations of Plagiarism Detection Systems State-of-the-art Plagiarism Detection approaches capably identify copy & paste and to some extent slightly modified Plagiarism. However, they cannot reliably identify strongly disguised Plagiarism forms, including paraphrases, translated Plagiarism, and idea Plagiarism, which are forms of Plagiarism more commonly found in scientific texts. This weakness of current systems results in a large fraction of today’s scientific Plagiarism going undetected.

  • comparative evaluation of text and citation based Plagiarism Detection approaches using guttenplag
    ACM IEEE Joint Conference on Digital Libraries, 2011
    Co-Authors: Bela Gipp, Norman Meuschke, Joeran Beel
    Abstract:

    Various approaches for Plagiarism Detection exist. All are based on more or less sophisticated text analysis methods such as string matching, fingerprinting or style comparison. In this paper a new approach called Citation-based Plagiarism Detection is evaluated using a doctoral thesis, in which a volunteer crowd-sourcing project called GuttenPlag identified substantial amounts of Plagiarism through careful manual inspection. This new approach is able to identify similar and plagiarized documents based on the citations used in the text. It is shown that citation-based Plagiarism Detection performs significantly better than text-based procedures in identifying strong paraphrasing, translation and some idea Plagiarism. Detection rates can be improved by combining citation-based with text-based Plagiarism Detection.

Paolo Rosso - One of the best experts on this subject based on the ideXlab platform.

  • algorithms and corpora for persian Plagiarism Detection
    Forum for Information Retrieval Evaluation, 2016
    Co-Authors: Habibollah Asghari, Paolo Rosso, Salar Mohtaj, Heshaam Faili, Omid S Fatemi, Martin Potthast
    Abstract:

    The task of Plagiarism Detection is to find passages of text-reuse in a suspicious document. This task is of increasing relevance, since scholars around the world take advantage of the fact that information about nearly any subject can be found on the World Wide Web by reusing existing text instead of writing their own. We organized the Persian PlagDet shared task at PAN 2016 in an effort to promote the comparative assessment of NLP techniques for Plagiarism Detection with a special focus on Plagiarism that appears in a Persian text corpus. The goal of this shared task is to bring together researchers and practitioners around the exciting topic of Plagiarism Detection and text-reuse Detection. We report on the outcome of the shared task, which divides into two subtasks: text alignment and corpus construction. In the first subtask, nine teams participated, whereas the best result achieved was a PlagDet score of 0.92. For the second subtask of corpus construction, five teams submitted a corpus, which were evaluated using the systems submitted for the first subtask. The results show that significant challenges remain in evaluating newly constructed corpora.

  • RuSSIR - Author Profiling and Plagiarism Detection
    Communications in Computer and Information Science, 2015
    Co-Authors: Paolo Rosso
    Abstract:

    In this paper we introduce the topics that we will cover in the RuSSIR 2014 course on Author Profiling and Plagiarism Detection (APPD). Author profiling distinguishes between classes of authors studying how language is shared by classes of people. This task helps in identifying profiling aspects such as gender, age, native language, or even personality type. In case of the Plagiarism Detection task we are not interested in studying how language is shared. On the contrary, given a document we are interested in investigating if the writing style changes in order to unveil text inconsistencies, i.e., unexpected irregularities through the document such as changes in vocabulary, style and text complexity. In fact, when it is not possible to retrieve the source document(s) where Plagiarism has been committed from, the intrinsic analysis of the suspicious document is the only way to find evidence of Plagiarism. The difficulty in retrieving the source of Plagiarism could be due to the fact that the documents are not available on the web or the plagiarised text fragments were obfuscated via paraphrasing or translation (in case the source document was in another language). In this overview, we also discuss the results of the shared tasks on author profiling (gender and age identification) and Plagiarism Detection that we help to organise at the PAN Lab on Uncovering Plagiarism, Authorship, and Social Software Misuse (http://pan.webis.de).

  • overview of the araplagdet pan fire2015 shared task on arabic Plagiarism Detection
    FIRE Workshops, 2015
    Co-Authors: Imene Bensalem, Paolo Rosso, Imene Boukhalfa, Lahsen Abouenour, Kareem Darwish, Salim Chikhi
    Abstract:

    is the first shared task that addresses the evaluation of Plagiarism Detection methods for Arabic texts. It has two sub- tasks, namely external Plagiarism Detection and intrinsic Plagiarism Detection. A total of 8 runs have been submitted and tested on the standardized corpora developed for the track. This overview paper describes these evaluation corpora, discusses the participants' methods, and highlights their building blocks that could be language dependent.

  • Plagiarism meets paraphrasing: Insights for the next generation in automatic Plagiarism Detection
    Computational Linguistics, 2013
    Co-Authors: Alberto Barrón-cedeño, Marta Vila, Maria Antònia Martí, Paolo Rosso
    Abstract:

    Although paraphrasing is the linguistic mechanism underlying many Plagiarism cases, little attention has been paid to its analysis in the framework of automatic Plagiarism Detection. Therefore, state-of-the-art Plagiarism detectors find it difficult to detect cases of paraphrase Plagiarism. In this article, we analyze the relationship between paraphrasing and Plagiarism, paying special attention to which paraphrase phenomena underlie acts of Plagiarism and which of them are detected by Plagiarism Detection systems. With this aim in mind, we created the P4P corpus, a new resource that uses a paraphrase typology to annotate a subset of the PAN-PC-10 corpus for automatic Plagiarism Detection. The results of the Second International Competition on Plagiarism Detection were analyzed in the light of this annotation. The presented experiments show that i more complex paraphrase phenomena and a high density of paraphrase mechanisms make Plagiarism Detection more difficult, ii lexical substitutions are the paraphrase mechanisms used the most when plagiarizing, and iii paraphrase mechanisms tend to shorten the plagiarized text. For the first time, the paraphrase mechanisms behind Plagiarism have been analyzed, providing critical insights for the improvement of automatic Plagiarism Detection systems.

  • A New Approach to Cross-Language Plagiarism Detection
    2013
    Co-Authors: Marc Franco-salvador, Parth Gupta, Paolo Rosso
    Abstract:

    Cross-language variant of automatic Plagiarism Detection tries to detect Plagiarism among documents across language pairs. In recent years a few approa- ches are proposed that use thesauri, alignment models or statistical dictionaries to deal with the similarity across languages. We propose a new approach to the cross- language Plagiarism Detection that makes use of a multilingual semantic network to generate knowledge graphs, obtaining a context model for each document which the other methods lack. To evaluate the proposed method, we use the Spanish-English and German-English partitions of the PAN-PC'11 corpus and compare our results with two state-of-the-art approaches. Experimental results indicate its potential to be a new alternative for similarity analysis in cross-language Plagiarism Detection.

Andreas Nurnberger - One of the best experts on this subject based on the ideXlab platform.

Vaclav Snasel - One of the best experts on this subject based on the ideXlab platform.