arXiv:2505.22637cs.LG2025-05被引 36

研究语言模型控制向量的不可靠性,发现方向不一致时效果反而适得其反。

Understanding (Un)Reliability of Steering Vectors in Language Models

  • 通过七种提示类型测试,发现激活差异方向不一致影响控制效果。
  • 训练集激活差值方向相似度越高,控制越有效。
  • 正负激活分离越好,模型越容易被精准引导。

控制向量是一种轻量级方法,通过在推理时向激活添加学习到的偏置来调控语言模型行为。尽管控制表现出良好潜力,但近期研究表明其在某些情况下可能不可靠甚至适得其反。本文研究了提示类型和激活差异几何结构对控制可靠性的影响。首先,我们发现所有七种提示类型均产生总体正向控制效应,但在样本间方差高,常出现与预期相反的效果;不同提示类型产生的控制向量方向差异显著(以余弦相似度衡量),无一种明显优于其他。其次,训练集中激活差异的余弦相似度越高,控制越有效。最后,正负激活在数据分布上越分离,模型越易被有效控制。结果表明,当目标行为未在特征空间中形成一致方向时,控制向量不可靠。

原文摘要 · Abstract (English)

Steering vectors are a lightweight method to control language model behavior by adding a learned bias to the activations at inference time. Although steering demonstrates promising performance, recent work shows that it can be unreliable or even counterproductive in some cases. This paper studies the influence of prompt types and the geometry of activation differences on steering reliability. First, we find that all seven prompt types used in our experiments produce a net positive steering effect, but exhibit high variance across samples, and often give an effect opposite of the desired one. No prompt type clearly outperforms the others, and yet the steering vectors resulting from the different prompt types often differ directionally (as measured by cosine similarity). Second, we show that higher cosine similarity between training set activation differences predicts more effective steering. Finally, we observe that datasets where positive and negative activations are better separated are more steerable. Our results suggest that vector steering is unreliable when the target behavior is not represented by a coherent direction.

语言模型控制向量可解释性提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。