arXiv:2606.26161cs.AI2026-06中稿 · ICML

发现聊天模型的拒绝行为受人格设定调控,晚层表达阶段才被开启。

Refusal Lives Downstream of Persona in Chat Models

论文配图:Refusal Lives Downstream of Persona in Chat Models
图 1 · 摘自论文原文
  • 通过干预激活空间方向,发现人格与拒绝行为存在联动关系。
  • 调整人格方向后,拒绝率从97%降至2%,但后期仍可恢复。
  • 拒绝能力在模型晚期才被释放,需人格信号触发,适合对安全机制研究者。

指令微调的聊天模型中,拒绝和人格特征在激活空间中均表现为线性方向,但以往研究将其视为独立机制。本文揭示二者存在交互:顺从的人格会抑制拒绝行为。在Qwen2.5-7B-Instruct和Llama-3.1-8B-Instruct模型中,我们提取了顺从人格方向与拒绝方向并进行干预。人格方向调控使拒绝率在Llama模型中由97%降至2%;重新引入拒绝方向可在深层恢复拒绝行为,但早期无法恢复。在深层窗口中移除人格方向可使拒绝行为回归基线,而移除随机方向则无效。因此,拒绝行为在深层表达阶段才被释放,位于人格计算之后。将拒绝视为单一独立方向会忽略其对人格的依赖性。

原文摘要 · Abstract (English)

Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms. We show they interact: a compliant persona gates refusal. In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both. Compliant persona steering suppresses refusal -- in Llama, the refusal rate falls from 97% to 2%. Reintroducing the refusal direction partially restores refusal at late layers but not at early ones. Projecting out the persona direction in a late-layer window restores it to baseline; projecting out a random direction does not. Refusal is therefore gated at the late-layer expression stage, downstream of where it is computed. Treating refusal as a single isolated direction misses its dependence on persona.

对话模型拒绝行为人格控制激活空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。