测试顶尖大模型在自动化越狱攻击下的安全性,发现仍可被低成本突破。
A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

- 用红队框架生成数十万次攻击,自动检测并验证成功案例。
- Opus 4.8有11.5%的有害意图被攻破,Fable 5为6.1%。
- 即使顶级模型也易被持续自动化攻击绕过,适合安全研究者参考。
我们评估了Anthropic开发的两个前沿大语言模型Fable 5和Opus 4.8,在7,826个涵盖十类危害的有害意图上,对四类自动化越狱攻击的对抗鲁棒性。通过HackAgent红队框架,生成数十万次对抗性尝试,并由三名裁判模型(多数投票)独立复核每一起成功案例。两个模型均抵御多数攻击,但残余漏洞远超整体评估所暗示:主要来自自适应迭代攻击,静态混淆几乎被完全破解。最强自适应搜索(树状攻击)使Opus 4.8在11.5%的意图上失效,而Fable 5最差情况为6.1%。聚合率不应被视为安全保证。即便在加固配置下,两模型仍分别产生1,620(Opus 4.8)和702(Fable 5)经面板确认的有害输出,覆盖所有危害类别,且由无专家参与的攻击模型在首次或第二次优化步骤内自动定位。合理结论是:即使最先进、经过充分测试的模型,在持续自动化压力下仍可被可靠攻破。
原文摘要 · Abstract (English)
We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundreds of thousands of adversarial attempts were generated and every apparent success was independently re-adjudicated by a panel of three judge models (majority vote). Both models resist the majority of attacks, but the residual surface is larger than aggregate framing suggests: it is dominated by adaptive iterative attacks, while static obfuscation is near-fully neutralised. The strongest adaptive search (tree-of-attacks) breaks Opus 4.8 on 11.5% of intents overall, whereas Fable 5 stays in the single digits (6.1% worst-case). Aggregate rates therefore should not be read as reassurance. Even in these hardened configurations, the two models produced 1 620 (Opus 4.8) and 702 (Fable 5) panel-confirmed harmful completions spanning every harm category, located automatically, cheaply, and within the first one or two refinement steps by an attacker model with no human expert in the loop. The reasonable conclusion is that even the best, most-tested frontier models remain reliably breakable under sustained automated pressure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。