评测顶尖AI模型安全防护能力,发现差距超百倍
AI Security Leaderboard: Methodology, Results and Minimal Standard

- 按最小安全标准测试四大模型的越狱防御能力
- 顶级模型拒绝对抗成本超1.4万美元,低端模型不足300美元
- 漏洞可修复,适合关注AI安全的开发者与监管者
AI安全排行榜是一项独立基准,从最不安全到最安全对前沿AI模型的防护能力进行排序。该榜单依据FAR.AI最小安全标准进行测试,此标准代表安全底线:未达标即意味着不具备最先进的安全性,但达标也不保证绝对安全。版本1.0涵盖化学、生物、辐射、核及爆炸(CBRNE)威胁和进攻性网络安全等严重滥用请求。本报告测试了四个领先模型在该标准下的通用越狱攻击表现,发现安全能力相差逾百倍。Claude Fable 5和GPT-5.6 Sol在所有攻击下均未被突破,估计越狱成本可能超过14,200美元;而Grok 4.5和Gemini 3.1 Pro存在数百个通用越狱路径,每个成本低于300美元,其中网络安全领域最弱的Grok越狱成本低至24美元。这些漏洞属于已有防御方案的已知攻击类型,且已在生产模型中部署。排行榜将随新模型发布持续更新,评估方法与最小标准也将定期修订以反映最新进展。排行榜网址:leaderboard.far.ai。
原文摘要 · Abstract (English)
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for security: meeting it does not guarantee a secure model, but failing to meet it guarantees a lack of state-of-the-art security. Version 1.0 covers severe misuse requests across chemical, biological, radiological, nuclear, and explosive (CBRNE) threats and offensive cybersecurity. In this report, we tested four leading models for universal jailbreaks in the context of this minimal standard, and found more than a hundredfold difference in security. Claude Fable 5 and GPT-5.6 Sol held against every attack we ran, with no universal jailbreak found; we estimate they would likely cost more than \$14,200 to jailbreak, if it is possible with this methodology at all. Meanwhile, we found hundreds of universal jailbreaks for Grok 4.5 and Gemini 3.1 Pro; each broke for under \$300, with universal jailbreaks in Grok's weakest domain, cybersecurity, accessible for as little as \$24. The gap is fixable: every weakness we found belongs to a known class of attack that already has a defense deployed in production models. The leaderboard will be updated on a rolling basis as new models are released, and the evaluation methodology and Minimal Standard will be periodically revised to take into account the latest capabilities and the state-of-the-art in safeguards. The leaderboard is available at leaderboard.far.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。