arXiv:2605.14240cs.LGcs.AI2026-05NAACL被引 3

测试多种文本检测方法在改写攻击下的鲁棒性,发现越强越脆弱。

Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods

论文配图:Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods
图 1 · 摘自论文原文
  • 对比三种检测方法及集成模型对改写攻击的抗性
  • 含Binoculars的集成模型效果最好但最易被攻破
  • 揭示性能与鲁棒性之间的根本矛盾,适合关注检测可信度的研究者

大型语言模型的兴起带来了抄袭和虚假信息传播等新挑战。随着绕过AI检测工具的出现,可靠的机器生成文本检测需求日益迫切。本文评估了多种机器生成文本检测方法在改写攻击下的鲁棒性,包括微调的RoBERTa、Binoculars以及文本特征分析方法,及其与随机森林分类器的集成。结果发现,包含Binoculars的集成模型表现最优,但在攻击下损失也最大。本研究揭示了当前先进检测技术中性能与鲁棒性之间的二元对立,挑战了对现有方法可靠性的普遍认知。

原文摘要 · Abstract (English)

The recent large-scale emergence of LLMs has left an open space for dealing with their consequences, such as plagiarism or the spread of false information on the Internet. Coupling this with the rise of AI detector bypassing tools, reliable machine-generated text detection is in increasingly high demand. We investigate the paraphrasing attack resilience of various machine-generated text detection methods, evaluating three approaches: fine-tuned RoBERTa, Binoculars, and text feature analysis, along with their ensembles using Random Forest classifiers. We discovered that Binoculars-inclusive ensembles yield the strongest results, but they also suffer the most significant losses during attacks. In this paper, we present the dichotomy of performance versus resilience in the world of AI text detection, which complicates the current perception of reliability among state-of-the-art techniques.

文本检测鲁棒性LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。