arXiv:2605.31113cs.CL2026-05被引 1

针对维基百科真实编辑场景,构建了更难检测的文本生成评测基准。

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

论文配图:TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices
图 1 · 摘自论文原文
  • 基于维基百科常见编辑任务设计多任务检测数据集
  • 现有模型在真实场景下准确率下降10%至40%
  • 任务特定训练可泛化到通用场景,反之则不行

自动检测机器生成文本(MGT)对维护用户生成内容平台(如维基百科)的知识完整性至关重要。现有检测基准主要关注通用生成任务(如“写一篇关于机器学习的文章”)。然而,编辑常使用大语言模型完成具体任务(如摘要)。这类任务特定的MGT因任务约束和上下文条件,更接近人类写作。本文发现,多种先进MGT检测器在真实维基百科编辑场景中表现显著下降。我们提出TSM-Bench,一个跨语言、多生成器、多任务的评测基准,用于评估检测器在常见真实编辑任务上的性能。结果表明:(i)平均检测准确率较以往基准下降10%–40%;(ii)存在泛化不对称性:在任务特定数据上微调可泛化至通用数据(甚至跨领域),但反向不成立。在通用MGT上微调的模型会过拟合于生成的表层特征。结果表明,与以往基准不同,多数检测器在真实场景下仍不可靠。TSM-Bench为未来模型的开发与评估提供了关键基础。

原文摘要 · Abstract (English)

Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such as Wikipedia. Existing detection benchmarks primarily focus on \textit{generic} text generation tasks (e.g., ``Write an article about machine learning.''). However, editors frequently employ LLMs for specific writing tasks (e.g., summarisation). These \textit{task-specific} MGT instances tend to resemble human-written text more closely due to their constrained task formulation and contextual conditioning. In this work, we show that a range of SOTA MGT detectors struggle to identify task-specific MGT reflecting real-world editing on Wikipedia. We introduce \textsc{TSM-Bench}, a multilingual, multi-generator, and \textit{multi-task} benchmark for evaluating MGT detectors on common, real-world Wikipedia editing tasks. Our findings demonstrate that (\textit{i}) average detection accuracy drops by 10--40\% compared to prior benchmarks, and (\textit{ii}) a generalisation asymmetry exists: fine-tuning on task-specific data enables generalisation to generic data -- even across domains -- but not vice versa. We demonstrate that models fine-tuned exclusively on generic MGT overfit to superficial artefacts of machine generation. Our results suggest that, in contrast to prior benchmarks, most detectors remain unreliable for automated detection in real-world contexts such as UGC platforms. \textsc{TSM-Bench} therefore provides a critical foundation for developing and evaluating future models.

文本检测大模型评估维基百科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。