越安全的模型越易被单次指令攻破,因自身判断力成漏洞。
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

- 用模型自己判断有害内容的能力反向诱导其输出有害回复。
- 30个开源模型中,安全能力越强,越容易被单次攻击破解。
- 首次揭示安全对齐与对抗攻击间的矛盾,适合安全研究者参考。
大型语言模型(LLMs)经过严格对齐以拒绝有害请求,这一过程使其内建了评估和识别不安全内容的能力。本文揭示,这种先进的安全意识反而引入致命漏洞。我们提出后验攻击(Posterior Attack),仅需一次提示即可绕过防护机制,诱使模型生成其内部分类器本应标记为不安全的回应。通过对30个开源模型(最大达35B参数)及前沿模型(如GPT-5、Claude 4.6)的广泛实证评估,发现安全判断能力越强的模型,越容易受到该攻击。我们形式化提出“安全悖论”,理论证明安全对齐的持续优化会自然增强后验脆弱性。通过强化学习干预验证因果关系:人为削弱模型的安全判断可使其免疫攻击,而增强判断则加剧漏洞。研究揭示当前对齐范式存在潜在缺陷,防御机制或需结构性重构。
原文摘要 · Abstract (English)
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability. We introduce Posterior Attack, a single-query jailbreak that bypasses guardrails by prompting the model to generate the exact harmful response its internal classifier would normally flag as unsafe. Through extensive empirical evaluation across 30 open-source LLMs (up to 35B parameters in size) and frontier models (e.g., GPT-5, Claude 4.6), we observe a striking phenomenon: models with superior safety-judgment capabilities are disproportionately more susceptible to this exploitation. To explain this, we formalize the Safety Paradox, analytically showing that monotonic improvements in safety alignment naturally amplify posterior vulnerability. Finally, we establish a causal link via reinforcement learning interventions, exemplifying that artificially degrading a model's safety judgment immunizes it against the attack, whereas enhancing judgment exacerbates the vulnerability. Our findings highlight potential flaws in current alignment paradigms, indicating that defense mechanisms may require further structural refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。