arXiv:2510.06805cs.CLcs.IR2025-10综述被引 3

构建大规模自动生成抄袭数据集,评估中文论文抄袭检测新挑战

Overview of the Plagiarism Detection Task at PAN 2025

  • 用 Llama、DeepSeek-R1、Mistral 生成抄袭文本,构建新数据集
  • 嵌入向量相似度方法达 0.8 召回率,但对旧数据集表现差
  • 揭示当前方法泛化能力弱,适合研究生成式抄袭检测的学者

PAN 2025 的生成式抄袭检测任务旨在自动识别科学文章中的自动生成抄袭内容,并将其与源文本对齐。我们利用 Llama、DeepSeek-R1 和 Mistral 三款大语言模型构建了一个全新的大规模自动生成抄袭数据集。本文概述了该数据集的构建过程,总结并比较了所有参赛者及四类基线方法的结果,并在 PAN 2015 的旧任务上评估了这些方法的表现,以分析其鲁棒性。结果显示,当前方法中基于嵌入向量的简单语义相似度方法在新数据集上达到最高 0.8 的召回率和 0.5 的精确率,但多数方法在 2015 年数据集上表现显著下降,表明现有方法缺乏泛化能力。

原文摘要 · Abstract (English)

The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective sources. We created a novel large-scale dataset of automatically generated plagiarism using three large language models: Llama, DeepSeek-R1, and Mistral. In this task overview paper, we outline the creation of this dataset, summarize and compare the results of all participants and four baselines, and evaluate the results on the last plagiarism detection task from PAN 2015 in order to interpret the robustness of the proposed approaches. We found that the current iteration does not invite a large variety of approaches as naive semantic similarity approaches based on embedding vectors provide promising results of up to 0.8 recall and 0.5 precision. In contrast, most of these approaches underperform significantly on the 2015 dataset, indicating a lack in generalizability.

抄袭检测大模型数据集评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。