arXiv:2604.15557cs.LGcs.CL2026-04被引 2

提出可预测转向向量效果的诊断工具,让模型干预更高效。

Predicting Where Steering Vectors Succeed

  • 用无训练的线性可访问性指标预测转向效果
  • 峰值指标相关性达0.86至0.91,准确指导层选择
  • 适用于大模型概念控制,实测验证有效

转向向量在某些概念和层上有效,但在其他情况下失败,从业者无法预先判断适用性。我们提出线性可访问性轮廓(LAP),一种针对每层的诊断方法,将对数透镜重用于预测转向向量有效性。核心指标 $A_{\mathrm{lin}}$ 通过模型的未嵌入矩阵作用于中间隐藏状态,无需训练。在五个模型(Pythia-2.8B 到 Llama-8B)上的24个受控二元概念族中,峰值 $A_{\mathrm{lin}}$ 对转向效果的预测相关性为 $ρ= +0.86$ 至 $+0.91$,对层选择的预测相关性为 $ρ= +0.63$ 至 $+0.92$。三阶段框架解释了何时差值法有效、何时需非线性方法、何时任何方法均无效。实体转向演示验证了全流程预测:在 Gemma-2-2B 和 OLMo-2-1B-Instruct 上,按 LAP 推荐层进行转向可成功改变生成结果,而标准中间层则无效果。

原文摘要 · Abstract (English)

Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running an intervention. We introduce the Linear Accessibility Profile (LAP), a per-layer diagnostic that repurposes the logit lens as a predictor of steering vector effectiveness. The key measure, $A_{\mathrm{lin}}$, applies the model's unembedding matrix to intermediate hidden states, requiring no training. Across 24 controlled binary concept families on five models (Pythia-2.8B to Llama-8B), peak $A_{\mathrm{lin}}$ predicts steering effectiveness at $ρ= +0.86$ to $+0.91$ and layer selection at $ρ= +0.63$ to $+0.92$. A three-regime framework explains when difference-of-means steering works, when nonlinear methods are needed, and when no method can work. An entity-steering demo confirms the prediction end-to-end: steering at the LAP-recommended layer redirects completions on Gemma-2-2B and OLMo-2-1B-Instruct, while the middle layer (the standard heuristic) has no effect on either model.

模型干预可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。