测试41个大模型安全性的新工具,发现多数模型极易被攻破。
LLM Robustness Leaderboard v1 --Technical report
- 用动态对抗优化自动红队测试,成功率100%
- 攻击难度差300倍,暴露模型间安全差异
- 可定位不同漏洞类型对应的突破方法,适合安全研究者
本技术报告配合PRISM Eval在巴黎人工智能行动峰会发布的大型语言模型鲁棒性排行榜。我们提出PRISM Eval行为诱导工具(BET),一个通过动态对抗优化实现自动化红队测试的AI系统,在41个前沿大模型中对其中37个实现了100%攻击成功率(ASR)。除二元成功指标外,我们引入细粒度鲁棒性度量,估算诱发有害行为所需的平均尝试次数,揭示攻击难度在模型间相差超过300倍,尽管所有模型均存在普遍脆弱性。我们还开展基础层级漏洞分析,识别针对特定危害类别的最有效越狱技术。通过与来自AI安全网络的可信第三方合作评估,展示了社区分布式鲁棒性测评的可行路径。
原文摘要 · Abstract (English)
This technical report accompanies the LLM robustness leaderboard published by PRISM Eval for the Paris AI Action Summit. We introduce PRISM Eval Behavior Elicitation Tool (BET), an AI system performing automated red-teaming through Dynamic Adversarial Optimization that achieves 100% Attack Success Rate (ASR) against 37 of 41 state-of-the-art LLMs. Beyond binary success metrics, we propose a fine-grained robustness metric estimating the average number of attempts required to elicit harmful behaviors, revealing that attack difficulty varies by over 300-fold across models despite universal vulnerability. We introduce primitive-level vulnerability analysis to identify which jailbreaking techniques are most effective for specific hazard categories. Our collaborative evaluation with trusted third parties from the AI Safety Network demonstrates practical pathways for distributed robustness assessment across the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。