用多语言回译绕过AI文本检测,暴露现有系统漏洞
ESPERANTO: Evaluating Synthesized Phrases to Enhance Robustness in AI Detection for Text Origination
- 通过多语言回译篡改AI生成文本,保持语义不变
- 9个检测器中真阳性率下降超40%,最高达67.8%降幅
- 开源72万条数据集,助力检测系统鲁棒性研究
尽管大语言模型在多个领域具有显著应用价值,但其也易被滥用于学术不端和虚假信息传播。为此,AI生成文本检测系统应运而生,但这些系统对文本篡改手段缺乏鲁棒性。本文提出一种基于回译的新型逃避检测技术:将AI生成文本经多语言翻译后再回译为英文,构建语义一致但可规避检测的篡改版本。实验评估涵盖6个开源与3个专有检测系统,结果显示,篡改后文本使多数检测器真阳性率(TPR)显著下降,最高降幅达67.8%。针对此问题,本文提出增强方法,使检测器在回译攻击下仅降低1.85%的TPR。同时构建包含8种大模型、72万条文本的公开数据集,覆盖多种领域与写作风格,支持检测系统性能评估。
原文摘要 · Abstract (English)
While large language models (LLMs) exhibit significant utility across various domains, they simultaneously are susceptible to exploitation for unethical purposes, including academic misconduct and dissemination of misinformation. Consequently, AI-generated text detection systems have emerged as a countermeasure. However, these detection mechanisms demonstrate vulnerability to evasion techniques and lack robustness against textual manipulations. This paper introduces back-translation as a novel technique for evading detection, underscoring the need to enhance the robustness of current detection systems. The proposed method involves translating AI-generated text through multiple languages before back-translating to English. We present a model that combines these back-translated texts to produce a manipulated version of the original AI-generated text. Our findings demonstrate that the manipulated text retains the original semantics while significantly reducing the true positive rate (TPR) of existing detection methods. We evaluate this technique on nine AI detectors, including six open-source and three proprietary systems, revealing their susceptibility to back-translation manipulation. In response to the identified shortcomings of existing AI text detectors, we present a countermeasure to improve the robustness against this form of manipulation. Our results indicate that the TPR of the proposed method declines by only 1.85% after back-translation manipulation. Furthermore, we build a large dataset of 720k texts using eight different LLMs. Our dataset contains both human-authored and LLM-generated texts in various domains and writing styles to assess the performance of our method and existing detectors. This dataset is publicly shared for the benefit of the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。