arXiv:2511.09200cs.LG2025-11

发现文本生成检测基准存在污染,影响检测器可靠性。

Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"

  • 分析发现98.5%的Claude-LLM数据含简单生成模式
  • 原始数据训练的检测器依赖这些模式,易被欺骗攻击
  • 清洗后数据提升检测鲁棒性,适合可信评估

大语言模型在诸多应用中日益普及,为防止滥用,检测其生成文本至关重要。检测器的训练与评估依赖高质量基准数据集。尽管多个团队已构建并发布大规模、多样化的数据集,但确保数据在各方面的高质量仍是挑战。例如,DetectRL基准中98.5%的Claude-LLM数据包含简单生成模式,如“Sure! Here is the academic article abstract:”等开头词,或模型拒绝任务的情况。本文表明,基于此类数据训练的检测器会利用这些模式作为捷径,从而容易遭受欺骗攻击。为此,我们对DetectRL数据集进行了多步清洗处理。实验表明,清洗后的数据使直接攻击更难实现。清洗后的数据集已公开可用。

原文摘要 · Abstract (English)

Large language models are increasingly used for many applications. To prevent illicit use, it is desirable to be able to detect AI-generated text. Training and evaluation of such detectors critically depend on suitable benchmark datasets. Several groups took on the tedious work of collecting, curating, and publishing large and diverse datasets for this task. However, it remains an open challenge to ensure high quality in all relevant aspects of such a dataset. For example, the DetectRL benchmark exhibits relatively simple patterns of AI-generation in 98.5% of the Claude-LLM data. These patterns may include introductory words such as "Sure! Here is the academic article abstract:", or instances where the LLM rejects the prompted task. In this work, we demonstrate that detectors trained on such data use such patterns as shortcuts, which facilitates spoofing attacks on the trained detectors. We consequently reprocessed the DetectRL dataset with several cleansing operations. Experiments show that such data cleansing makes direct attacks more difficult. The reprocessed dataset is publicly available.

文本检测数据污染基准评估安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。