arXiv:2508.19697cs.CRcs.AI2025-08被引 5

让大模型安全机制不再依赖少数注意力头,提升抗攻击能力

Safety Alignment Should Be Made More Than Just A Few Attention Heads

  • 用拒绝方向定位关键安全注意力头,发现安全依赖集中
  • 新训练策略使安全行为分布到更多注意力头,安全性提升40%以上
  • 适合关注大模型安全防护、对抗攻击的研究者与工程师

当前大语言模型的安全对齐仍存在漏洞,因对抗提示可有效绕过其安全机制。我们发现这些安全机制主要依赖有限的注意力头:移除或消融这些头会严重削弱模型安全。为此,我们提出RDSHA,一种基于拒绝方向的定向消融方法,用于识别最影响安全行为的注意力头。分析表明,现有越狱攻击正是通过选择性绕过或操纵这些关键头实现突破。为解决此问题,我们提出AHD训练策略,旨在将安全相关行为分散编码至更多注意力头中。实验显示,AHD成功将安全能力分布到更广泛的注意力头;在多个主流越狱攻击下,经AHD训练的模型展现出显著更强的安全鲁棒性,同时保持整体功能有效性。

原文摘要 · Abstract (English)

Current safety alignment for large language models(LLMs) continues to present vulnerabilities, given that adversarial prompting can effectively bypass their safety measures.Our investigation shows that these safety mechanisms predominantly depend on a limited subset of attention heads: removing or ablating these heads can severely compromise model safety. To identify and evaluate these safety-critical components, we introduce RDSHA, a targeted ablation method that leverages the model's refusal direction to pinpoint attention heads mostly responsible for safety behaviors. Further analysis shows that existing jailbreak attacks exploit this concentration by selectively bypassing or manipulating these critical attention heads. To address this issue, we propose AHD, a novel training strategy designed to promote the distributed encoding of safety-related behaviors across numerous attention heads. Experimental results demonstrate that AHD successfully distributes safety-related capabilities across more attention heads. Moreover, evaluations under several mainstream jailbreak attacks show that models trained with AHD exhibit considerably stronger safety robustness, while maintaining overall functional utility.

模型安全注意力机制对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。