普通系统提示变化可让越狱攻击成功率飙升56个百分点,暴露安全评估漏洞。
The Fragility of Jailbreak Robustness Across Operational States

- 通过改变非安全相关系统提示,研究操作状态对越狱攻击的影响。
- 在7个对齐模型上,越狱成功率在不同状态下相差高达56个百分点。
- 揭示隐藏表示中拒绝轴的差异可预测越狱成功,适合安全评测研究人员。
现有越狱评估通常用单一攻击成功率(ASR)衡量鲁棒性,且在默认配置(原始状态)下进行。然而,用户与大模型的交互可能引发多种非原始操作状态。本文发现,越狱鲁棒性对操作状态变化极为脆弱:即使攻击方式不变,仅更改一个未设计用于影响安全性的普通系统提示,就能显著改变攻击成功率。我们在7个对齐模型和3种典型越狱攻击上系统研究此现象,观察到原始状态与非原始状态间ASR存在显著差异。在某一案例中,仅因操作状态变化,ASR从2%上升至58%,增幅达56个百分点。令人惊讶的是,这种提升甚至出现在原本在原始状态下设计优化的攻击中。我们进一步发现,状态依赖的鲁棒性变化与拒绝相关轴上的隐藏表示差异密切相关,该轴上的投影能强预测越狱结果。结果表明,单一原始状态评估无法全面刻画越狱鲁棒性,亟需拓展至非原始操作状态的评估。
原文摘要 · Abstract (English)
Existing jailbreak evaluations typically characterize robustness using a single attack success rate (ASR) measured in a default configuration (the vanilla state). However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates. We systematically investigate this phenomenon across seven aligned models and three representative jailbreak attacks, observing substantial differences in ASR between vanilla and non-vanilla operational states. In one case, ASR increases by up to 56 percentage points (2% to 58%) solely due to a change in operational state. Remarkably, these increases occur even for attacks originally designed and optimized under vanilla-state evaluation. We further show that state-dependent robustness variation is systematically associated with differences in hidden representations along a refusal-related axis, and that projections onto this axis strongly predict jailbreak outcomes. Our results show that a single vanilla-state evaluation may not fully characterize jailbreak robustness, motivating evaluations that also examine how robustness changes across non-vanilla operational states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。