检测AI润色文本时,现有工具误判率高且无法区分程度。
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
- 构建14.7K样本数据集,测试不同程度AI润色的检测效果。
- 多数检测器将轻微润色文本误判为完全AI生成。
- 适合关注内容安全与检测公平性的研究人员参考。
大型语言模型(LLMs)广泛应用引发对生成文本检测的担忧,但一个被忽视的问题是:人类撰写的内容经由AI工具进行细微润色。这带来关键问题:轻微润色是否应归类为AI生成?此类误判可能导致错误的抄袭指控,并夸大在线内容中AI占比。本研究通过自建的AI-Polished-Text Evaluation(APT-Eval)数据集,系统评估了十二种先进文本检测器,该数据集包含14.7K个在不同程度上受AI影响的样本。结果表明,检测器常将极轻微润色文本误标为AI生成,难以区分不同层次的AI参与度,并对较老及较小模型表现出偏见。这些局限凸显了开发更细致检测方法的紧迫性。
原文摘要 · Abstract (English)
The growing use of large language models (LLMs) for text generation has led to widespread concerns about AI-generated content detection. However, an overlooked challenge is AI-polished text, where human-written content undergoes subtle refinements using AI tools. This raises a critical question: should minimally polished text be classified as AI-generated? Such classification can lead to false plagiarism accusations and misleading claims about AI prevalence in online content. In this study, we systematically evaluate twelve state-of-the-art AI-text detectors using our AI-Polished-Text Evaluation (APT-Eval) dataset, which contains 14.7K samples refined at varying AI-involvement levels. Our findings reveal that detectors frequently flag even minimally polished text as AI-generated, struggle to differentiate between degrees of AI involvement, and exhibit biases against older and smaller models. These limitations highlight the urgent need for more nuanced detection methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。