模型的默认行为会干扰指令控制,而非被控行为本身决定影响范围。
Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

- 将指令方向作为扰动,模型会向其偏好行为靠拢。
- 24种行为中仅20种受实际影响,且主要偏向拒绝、奉承和诗意表达。
- 无论操控什么行为,干扰模式基本一致,适合研究模型内在偏好者阅读。
激活控制(activation steering)承诺对语言模型行为实现模块化调控:某一行为(如礼貌)对应激活空间中的一个方向,添加该方向应仅切换此行为而保持其余不变。然而事实并非如此。我们探究了哪些其他行为会被影响以及程度如何,发现决定因素是模型本身而非被操控的行为。一次控制会使模型趋向于其固有的小范围偏好行为,主要包括拒绝、奉承和诗意表达,且这一偏好集合在不同控制任务中基本一致。在10个指令微调模型上,对24种行为进行测试,所有结果均通过语言模型判别器从生成文本中读取,而非依赖探测器。重要的是,尽管所有24种行为均可线性解码,但仅有20种真正改变生成内容。第一,一个无行为语义、仅大小匹配真实控制向量的方向,仍能引发与真实控制相同的干扰顺序,却无法产生需要特定方向的行为;第二,大多数干扰具有单向性,例如控制粗俗会引发毒性,但控制毒性不会影响粗俗;第三,当某行为完全被排除时,其余行为的几何关系几乎无法解释其参与的干扰。该结论在全部10个模型上成立,且低于100亿参数的模型最明显,大型模型则减弱。将控制视为模型固定终点的扰动表明,仅解耦行为方向不足以实现真正的模块化控制。
原文摘要 · Abstract (English)
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。