arXiv:2607.13565cs.CRcs.AI2026-07

通过结构化分布外扰动,突破当前最先进检测器的防御屏障。

UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors

  • 利用文本分布外偏移策略,避开对抗性微调的检测
  • 攻击成功率提升约50倍,且保持文本自然度
  • 适合研究检测系统脆弱性与对抗样本攻防的学者

我们研究了哪些语言模型逃避攻击能在最先进的对抗性微调后仍有效,开发出能登上ELOQUENT 2026 Voight-Kampff排行榜前五名的策略。尽管对抗性微调轻易关闭了2025年获胜的逃避方法,我们发现检测器存在根本性不对称漏洞:将生成文本推离检测器训练数据分布可稳定绕过检测,而将其拉入分布(如模仿人类训练数据)则完全无效。基于此,我们提出两种新型分布外攻击方法——跨年代注册攻击与现代主义意识流形式。这两种策略均能轻松规避对抗性闭合,使欺骗率比此前方法提高约50倍,同时保持文本自然性。此外,实验表明,部署方常见的应对措施(在训练数据中加入时期文体)无法弥补这一漏洞。研究结果表明,包括对抗性微调在内的检测模型,在面对结构性分布外偏移时仍存在持续性脆弱,这一机制直接支撑了我们在竞赛中的领先表现。

原文摘要 · Abstract (English)

We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the top 5 positions on the ELOQUENT 2026 Voight-Kampff leaderboard. While adversarial fine-tuning trivially closes the 2025 winning evasion recipes, we uncover a fundamental asymmetry in detector vulnerability: pushing generated text out of the detector's training distribution reliably defeats adversarial detection, whereas pulling it into the distribution (e.g., mimicking human training data) fails completely. Exploiting this, we introduce two novel out-of-distribution attack families - cross-decade register attacks and modernist stream-of-consciousness form. Both strategies easily bypass adversarial closure, achieving up to approximately 50x higher fool rates than previous methods while preserving naturalness. Furthermore, experiments show that the obvious deployer countermeasure (augmenting training data with period prose) fails to close the vulnerability. Our findings show that the tested detector families, including adversarially fine-tuned ones, exhibit persistent vulnerabilities under structural out-of-distribution shifts, a mechanism that directly powers our leading competition performance.

文本检测对抗攻击分布外大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。