越先进的模型,越难被越狱攻击拖垮性能。
Jailbroken Frontier Models Retain Their Capabilities

- 越强大的模型,越能扛住越狱攻击带来的性能损耗。
- 顶级模型越狱后性能仅下降7.7%,远低于普通模型的33.1%。
- 推理类任务受攻击影响更大,但最强越狱手段几乎无损绕过安全检测。
随着语言模型防护机制日益强化,攻击者被迫采用更复杂的越狱方法。以往研究发现,这种复杂性会带来‘越狱税’,降低目标模型的任务表现。我们发现该代价与模型能力呈反比:最先进的越狱攻击对顶尖模型几乎不造成性能损失。在涵盖从 Haiku 4.5 到 Opus 4.6 的五种 Claude 模型上评估 28 种越狱方法,结果显示:Haiku 4.5 在越狱后平均性能下降 33.1%,而 Opus 4.6 在最大思考努力下仅下降 7.7%。此外,所有模型中,推理类任务的退化程度显著高于知识回忆类任务。当前最强的边界点越狱(Boundary Point Jailbreaking)可实现近乎完美的分类器规避,且在受保护模型上性能几乎无损。因此,我们建议:前沿模型的安全性评估不应依赖于越狱导致显著能力下降这一假设。
原文摘要 · Abstract (English)
As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this complexity imposes a "jailbreak tax" that degrades the target model's task performance. We show that this tax scales inversely with model capability and that the most advanced jailbreaks effectively yield no reduction in model capabilities. Evaluating 28 jailbreaks on five benchmarks across Claude models ranging in capability from Haiku 4.5 to Opus 4.6, we find Haiku 4.5 loses an average of 33.1% on benchmark performance when jailbroken, while Opus 4.6 at max thinking effort loses only 7.7%. We also observe that across all models, reasoning-heavy tasks display considerably more degradation than knowledge-recall tasks. Finally, Boundary Point Jailbreaking, currently the strongest jailbreak against deployed classifiers, achieves near-perfect classifier evasion with near-zero degradation across safeguarded models. We recommend that safety cases for frontier models should not rely on a meaningful capability degradation from jailbreaks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。