通过描述长度增益检测大模型生成文本的源引用,有效应对改写与多源合成挑战。
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking

- 基于概率预测的描述长度理论,比较文本在有无候选来源时的编码长度差异。
- 在PAN 2025和PAN 2026基准上,准确率超92%,多源检索召回率达96%。
- 无需训练、可分解到词粒度,适合检测深度改写后的抄袭行为。
大语言模型对学术诚信构成威胁,但生成式抄袭检测仍属未解难题。以往方法关注AI生成痕迹而非源内容复用,相似性方法在深度改写或多源合成后失效。本文提出源条件描述长度增益(SCDG),基于概率预测中附加信息可压缩目标序列编码长度的原理,对比冻结语言模型在有无候选来源条件下对可疑文档的描述长度,得到逐标记的对数似然增益,量化来源提供的增量预测证据。在CLEF PAN基准上,基于PAN 2025的成对任务中,SCDG实现0.92精确率、0.97召回率、0.94 F1,优于所有基线;在PAN 2026多源检索任务中,nDCG@10达0.83,Recall@100为0.96。在同主题同事件的Multi-News测试中,校准后的增益分布分类器仅0.125%的样本被误判为源复用,表明其对主题重叠具有鲁棒性。结果证明SCDG是应对复杂变换下的源特定内容复用的统一、可分解信号。
原文摘要 · Abstract (English)
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, which may be permissible, rather than source reuse, while similarity-based methods struggle after extensive rewriting and multi-source synthesis. Motivated by the description-length view of probabilistic prediction, in which relevant side information can reduce a target sequence's code length, we introduce Source-Conditioned Description-Length Gain (SCDG), a directional, training-free framework that contrasts a frozen language model's description length of a suspicious document $P$ with and without a candidate source $S$. This contrast yields token-level log-likelihood gains that measure the incremental predictive evidence supplied by $S$. We evaluate SCDG on the PAN at CLEF benchmarks for generative plagiarism. On a PAN 2025-derived pairwise benchmark, SCDG achieves 0.92 Precision, 0.97 Recall, and 0.94 F1, outperforming all baselines; on PAN 2026's multi-source retrieval task, it reaches 0.83 nDCG@10 and 0.96 Recall@100, surpassing all baselines. On a same-topic, same-event Multi-News test, the calibrated gain-distribution SCDG classifier predicts source reuse for only $0.125\%$ of pairs, supporting robustness to topical overlap under this evaluation protocol. These results establish SCDG as a unified and token-decomposable signal for source-specific content reuse under extensive transformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。