揭示语言模型控制向量不可靠的原因及改进方向
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
- 通过几何相似性分析预测控制向量可靠性
- 训练数据中正负激活分离度越高,控制越稳定
- 适合关注模型可控性与非线性表征的研究者
控制向量是一种轻量级方法,通过在推理时向激活添加学习到的偏置来调控语言模型行为。尽管平均效果良好,但不同样本间的控制效果差异大,许多目标行为控制不可靠。本文研究发现:训练激活差异的余弦相似度越高,控制越可靠;当正负样本在控制方向上分离更明显时,控制更稳定;不同提示变体训练出的控制向量方向各异,但性能相似且跨数据集表现相关。结果表明,当潜在行为表示无法被线性控制方向有效近似时,控制向量便不可靠。这些发现为诊断控制不可靠性提供实用工具,并推动开发考虑非线性行为表示的更鲁棒控制方法。
原文摘要 · Abstract (English)
Steering vectors are a lightweight method for controlling language model behavior by adding a learned bias to the activations at inference time. Although effective on average, steering effect sizes vary across samples and are unreliable for many target behaviors. In my thesis, I investigate why steering reliability differs across behaviors and how it is impacted by steering vector training data. First, I find that higher cosine similarity between training activation differences predicts more reliable steering. Second, I observe that behavior datasets where positive and negative activations are better separated along the steering direction are more reliably steerable. Finally, steering vectors trained on different prompt variations are directionally distinct, yet perform similarly well and exhibit correlated efficacy across datasets. My findings suggest that steering vectors are unreliable when the latent target behavior representation is not effectively approximated by the linear steering direction. Taken together, these insights offer a practical diagnostic for steering unreliability and motivate the development of more robust steering methods that explicitly account for non-linear latent behavior representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。