arXiv:2602.02132cs.CL2026-02被引 13

大模型拒答行为不止一种方向,不同拒绝方式对应不同空间方向。

There Is More to Refusal in Large Language Models than a Single Direction

  • 发现拒绝行为在激活空间中有多条几何独立的方向
  • 任一拒绝方向线性操控均产生相似的拒绝与过度拒绝权衡
  • 适合研究模型安全机制或可控生成的开发者参考

以往研究认为大语言模型的拒绝行为由单一激活空间方向决定,可实现有效调控与消融。我们发现该观点不完整:在十一类拒绝与非合规行为(包括安全、请求不完整或无支持、拟人化、过度拒绝等)中,这些拒绝行为对应激活空间中几何上互异的方向。尽管方向多样,沿任意拒绝相关方向进行线性操控,都会产生几乎相同的拒绝与过度拒绝之间的权衡,如同一个共享的一维控制旋钮。不同方向的主要影响并非是否拒绝,而在于如何拒绝。

原文摘要 · Abstract (English)

Prior work argues that refusal in large language models is mediated by a single activation-space direction, enabling effective steering and ablation. We show that this account is incomplete. Across eleven categories of refusal and non-compliance, including safety, incomplete or unsupported requests, anthropomorphization, and over-refusal, we find that these refusal behaviors correspond to geometrically distinct directions in activation space. Yet despite this diversity, linear steering along any refusal-related direction produces nearly identical refusal to over-refusal trade-offs, acting as a shared one-dimensional control knob. The primary effect of different directions is not whether the model refuses, but how it refuses.

拒绝行为激活空间可控生成模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。