arXiv:2603.24511cs.LGcs.AI2026-03被引 11

AI自动发现新型攻击方法,突破大模型防御极限。

Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

  • 用AI代理在自研循环中自动探索攻击算法。
  • 对GPT-OSS-Safeguard-20B攻击成功率达80%,超旧方法50%以上。
  • 适用于对抗训练模型,适合安全研究者与防御评估者参考。

我们证明,AI代理能够自主发现针对大语言模型的新型对抗攻击算法,在白盒越狱和提示注入测试中达到前沿水平。通过部署前沿代理(如Claude Code和Codex),在有限算力预算下,结合30余种已有方法库与固定评估脚本,构建自动化研究流程。该流程成功攻破OpenAI的GPT-OSS-Safeguard-20B,在CBRN查询上实现最高80%的攻击成功率,远超现有方法<50%的表现;同时对Meta的抗攻击模型SecAlign-70B实现100%攻击成功率,优于此前最佳自动化方法的82%。值得注意的是,这些攻击方法是在无关代理模型上针对随机目标词强制任务生成,却能直接泛化至对抗训练模型的提示注入攻击。我们还追溯了自动研究过程中方法的演化路径,分析代理策略与失败模式。对抗机器学习长期强调防御需针对定制化攻击评估,而自动研究实现了这一原则的自动化,我们认为这应成为未来防御评估的最低标准。

原文摘要 · Abstract (English)

We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations. We deploy frontier agents, such as Claude Code and Codex, in an autoresearch loop with access to a library of 30+ prior methods and an evaluation script with a fixed compute budget. We show this pipeline to be effective in jailbreaking OpenAI's GPT-OSS-Safeguard-20B and in prompt injections against Meta-SecAlign-70B, an adversarially robust model. For GPT-OSS-Safeguard, the best agent-discovered method achieves up to 80\% attack success rate on CBRN queries, compared to <50\% for existing methods. For SecAlign, it achieves 100\% ASR, while the best prior automated methods only achieve 82\%. Notably, in our setting, attack methods are developed on unrelated surrogate models for a pure random-target token-forcing task, yet generalize directly to prompt injection on the adversarially trained model. Finally, we trace the lineage of methods developed during autoresearch, characterizing the agents' strategies and failure modes. Adversarial ML has long held that defenses must be evaluated against attacks tailored to them; autoresearch automates this principle, and we argue it should be the minimum bar for defense evaluation going forward.

对抗攻击AI代理大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。