语言模型拒答时,答案仍可从内部状态恢复,但重置拒答却难得多。
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

- 通过精准控制实验,发现拒答与回答的激活轨迹对称
- 释放被隐藏答案只需单点干预,但重新拒答需多点操作
- 拒绝行为非对称开关,探测恢复不能代表可控性
当语言模型拒绝回答时,正确答案是否已被彻底抹除,或仅在输出层被抑制尚不明确。我们通过受控屏蔽设置,实现了问答与拒答轨迹的精确匹配,用于双向激活补丁干预。研究揭示了在相同因果干预下存在因果不对称性,称为‘破缺对称’。即使模型生成了清晰的拒答,正确答案仍可从隐藏状态中线性恢复。释放该隐藏答案是高度局部的操作,仅需单位置补丁即可实现。而反向操作——重新施加抑制则不具同等局部性,需在多个位置进行广泛干预,且重构连贯拒答序列更为困难。我们进一步证明,平均的答案到拒答位移向量虽能刻画两状态间的几何差异,但无法作为可靠、可逆的线性控制开关。综上,拒答并非简单的对称开关。对安全与审计而言,这意味着探测可恢复性可能夸大真实行为控制能力,定位拒答相关方向并不保证能可靠引导模型从回答转向一致拒答。
原文摘要 · Abstract (English)
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we term broken symmetry. Even when a model generates a clean refusal, the correct answer remains linearly recoverable from its hidden states. Furthermore, releasing this withheld answer is a highly local operation, requiring only a single-position patch. Conversely, the reverse operation is not equally local: reimposing suppression requires broader interventions across multiple positions, and assembling a coherent refusal sequence is more difficult still. We further demonstrate that while an average answer-to-refusal displacement vector marks the geometric difference between these states, it fails to act as a reliable, reversible linear control toggle between behaviours. Taken together, our findings show that refusal does not function as a simple symmetric switch. For safety and auditing, this implies that probe recoverability can overestimate true behavioural control, and locating refusal-relevant directions does not reliably grant the ability to steer a model from answering to coherent refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。