研究发现拒绝行为的几何结构源于训练,多样化的拒绝开头可增强模型抗攻击能力。
Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

- 通过分析训练动态,发现拒绝方向由拒绝完成损失驱动。
- 重复的拒绝起始导致梯度集中于低维子空间,使拒绝易被移除。
- 引入多样化拒绝前缀可提升梯度稳定秩,抵抗向量消融攻击。
拒绝训练通过让模型拒绝不安全请求来保护AI免受越狱攻击,降低滥用风险。近期研究发现,对齐语言模型中的拒绝行为可通过单一激活方向或跨有害提示共享的低维拒绝子空间实现:消除这些方向可抑制拒绝行为,同时基本保留其他模型能力。然而,为何安全关键特征在多种模型中出现并呈现集中、低维结构仍不清楚。以OLMo-2-0425-1B-Instruct为例,我们发现拒绝几何结构反映拒绝训练过程:拒绝完成任务的一次性损失所引发的激活更新解释了最终的拒绝方向与子空间。通过分析不同拒绝数据集的训练动态,我们揭示其脆弱性与重复拒绝起始相关,后者又与梯度和拒绝特征在低维子空间的集中有关。在冻结模型分析与受控合成微调实验中,我们发现存在一个加固机制:多样化拒绝起始可提高梯度与激活变化的稳定秩,使拒绝更难通过向量消融攻击移除。
原文摘要 · Abstract (English)
Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge and concentrated, low-dimensional structure. In a case study of OLMo-2-0425-1B-Instruct we find that the refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace. We study refusal directions through the training dynamics across refusal datasets and reveal that their brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace. Across frozen-model analyses and controlled synthetic fine-tuning, we find evidence of a hardening lever: diverse refusal starts can raise stable ranks of gradients and activation changes, making refusals harder to remove with a vector ablation attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。