攻击不摧毁安全特征,而是针对性抑制特定注意力头。
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

- 发现攻击仅抑制早期层的对抗性注意力头,而中期层的安全头保持活跃。
- 抑制少数对抗性头即可诱发越狱行为,移除安全头则削弱防御能力。
- 无需训练,读取持续激活信号即可实现强鲁棒检测,适合模型安全研究者。
越狱攻击可绕过大语言模型的安全对齐,但其机制尚不明确。我们发现攻击并非全面消除安全特征,而是选择性抑制特定注意力头。识别出两类功能分化头:集中在早期层的对抗性受损头(ACHs)在攻击下被抑制,而位于中层的安全对齐头(SAHs)在攻击成功时仍维持强激活。消融实验表明ACHs具有因果作用,而SAHs贡献于鲁棒激活:抑制少量ACHs足以使正常拒绝输入产生越狱行为;移除SAHs则显著削弱中层安全激活。词元级归因显示,ACH抑制由攻击模板词元驱动,解释了为何攻击能通过抑制ACH绕过拒绝决策,同时留下的安全信号由SAHs维持——这一现象称为鲁棒有害特征。为验证其实际意义,我们证明仅读取这些持续激活信号,无需训练即可实现与强对抗鲁棒性相当的聚合检测性能。
原文摘要 · Abstract (English)
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate safety features, but instead selectively suppress specific attention heads. We identify two functionally differentiated types: Adversarially Compromised Heads (ACHs) concentrated in early layers, which are suppressed under attacks, and Safety-Aligned Heads (SAHs) in mid-layers, which maintain robust activations even when attacks succeed. Ablation studies support the causal role of ACHs and the contribution of SAHs to robust activations: suppressing a small number of ACHs is sufficient to induce jailbreak-like behavior on normally refused inputs, while removing SAHs substantially weakens mid-layer safety activations. Token-level attribution further shows that ACH suppression is driven specifically by attack-template tokens, providing a mechanistic account of why attacks can bypass refusal decisions through ACH suppression while leaving internal safety signals sustained by SAHs -- a phenomenon we term Robust Harmful Features. To validate the practical significance of this robustness, we show that simply reading these persistent activations -- without any training -- yields competitive aggregate detection performance with strong adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。