用防御者视角评估越狱攻击,提升模型安全能力。
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

- 以安全训练效果为标准,重新评估越狱攻击价值。
- 仅需少量测试即可准确估算攻击贡献度,选出最优攻击组合。
- 相比传统方法,直接优化攻击组合能更好提升模型安全性。
大型语言模型的越狱攻击通常通过攻击成功率(ASR)等攻击者中心指标评估,但破坏模型并不等于有助于提升安全性。本文提出一种防御者中心视角,将越狱攻击视为安全训练的红队数据,评估其带来的下游安全改进。基于此,提出A-MESS框架,通过黑箱子集效用观察,实现攻击的归因与选择。A-MESS计算基于Shapley值的AttackSHAP得分,量化单个攻击的边际贡献,并在用户设定预算下,通过贪心或代理优化选择紧凑攻击子集。在可控效用场景和真实大模型安全设置中,我们发现ASR排名与防御者效用弱相关;AttackSHAP可仅用少量效用查询准确估计;直接优化攻击子集的性能优于仅按攻击者视角或仅归因的选择方式。结果表明,应将越狱攻击视为提升安全性的资源,而不仅是突破模型的工具。
原文摘要 · Abstract (English)
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting-agnostic framework for attributing and selecting jailbreak attacks from black-box subset utility observations. A-MESS estimates AttackSHAP, a Shapley-based score that attributes marginal utility to individual attacks and selects compact attack subsets under user-specified budgets via greedy or surrogate-based optimization. Across controlled utility landscapes and real LLM safety settings, we find that ASR rankings are weakly aligned with defender-centric utility, that AttackSHAP can be estimated accurately with limited utility queries, and that directly optimizing subsets yields stronger safety utility than attacker-centric or attribution-only selection. These results suggest evaluating jailbreak attacks as resources for improving safety, not only as tools for breaking models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。