arXiv:2606.21077cs.CRcs.CL2026-06

用少量词替换突破毒性过滤,让恶意提示绕过检测

OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization

论文配图:OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization
图 1 · 摘自论文原文
  • 通过替换少数关键词,实现毒性与恶意意图解耦
  • 在4个GPT模型上平均攻击成功率从7.0%提升至84.0%
  • 适合安全审计人员和模型防护开发者参考

生产级大模型主要依赖基于毒性的内容过滤机制,假设有害意图与表面毒性表述相关。我们证明这一假设本质脆弱:仅通过替换五个词,即可实现表面毒性与对抗性意图的解耦。本文提出OTTER(Obfuscated Toxicity-Evading Token Evolution for Rewriting),一种仅需标准API访问的黑盒红队框架,直接针对工业安全审计的实际约束。在457个AdvBench提示上对四个GPT模型进行评估,OTTER将平均攻击成功率(ASR)从7.0%提升至84.0%。我们还首次提供毒性与绕过关系的量化分析及按类别分解结果,为生产部署中的分类器加固提供可操作建议。

原文摘要 · Abstract (English)

Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assumption is fundamentally brittle: surface toxicity and adversarial intent can be decoupled by replacing as few as five tokens. We present OTTER (Obfuscated Toxicity-Evading Token Evolution for Rewriting), a black-box red-teaming framework requiring only standard API access, directly targeting the practical constraints of industry security audits. Evaluated on 457 AdvBench prompts across four GPT models, OTTER raises average ASR from 7.0% to 84.0%. We further provide the first quantitative analysis of the toxicity--bypass relationship and a per-category breakdown, translating our findings into actionable recommendations for classifier hardening in production deployments.

红队攻击提示劫持模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。