arXiv:2607.00572cs.AIcs.CR2026-07

通过耦合有害性与拒绝响应方向,提升大模型安全对齐的鲁棒性。

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

论文配图:HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
图 1 · 摘自论文原文
  • 在提示和生成阶段同时耦合有害性与拒绝方向,增强安全机制
  • 在五个模型家族、两种规模下均有效,且不降低通用能力
  • 无需特定架构调优,适用于多种训练和推理阶段的安全方法

理解大语言模型如何内部分别表征有害性和拒绝行为,对诊断对齐漏洞至关重要,可解释越狱攻击为何成功并指导稳健对齐策略的设计。已有研究发现,对齐后的模型在提示侧词元位置的残差流中将有害性和拒绝编码为可分离的方向。我们发现,越狱攻击通过在生成任何词元前抑制拒绝或有害性方向而成功,且不同攻击类型占据有害性-拒绝平面上不同的区域。将分析扩展到响应词元位置后,我们发现模型在生成有害内容时仍能识别其危害性,即使在提示侧未能识别。基于此,我们提出HARC(有害性与拒绝耦合)微调方法,跨提示和响应位置配对两个方向。由于干预局限于有害性-拒绝子空间,其余残差流保持不变,不损害通用能力,也避免过度拒绝。在广泛实验中,HARC在六种基线方法中实现了最强的鲁棒性-能力-可用性平衡。在五种模型家族和两种规模下,提示与响应位置的有害性与拒绝方向均能迁移,无需架构特异性调优。

原文摘要 · Abstract (English)

Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.

安全对齐越狱攻击微调方法大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。