arXiv:2607.25270cs.CLcs.AI2026-07中稿 · EMNLP被引 2

发现模型即将执行行为时的激活状态才是有效引导信号来源。

Where Steering Signals Come From: Activation Source Selection in Activation Steering

论文配图:Where Steering Signals Come From: Activation Source Selection in Activation Steering
图 1 · 摘自论文原文
  • 通过分析模型执行边界状态,定位最有效的引导信号
  • 尾部减法技术使引导信号更稳定清晰,提升成功率
  • 适合研究可控生成与模型干预的从业者

激活引导在推理时通过向隐藏状态添加向量或特征来控制语言模型,但其上游信号来源常被忽视。本文将此问题定义为激活源选择:即用于提取构建引导信号的隐藏状态的上下文与读取策略组合。在三个指令微调模型和四种引导任务上,固定下游干预的情况下,仅改变激活源就显著影响引导效果。研究发现,有效引导并非取决于目标行为是否出现在源文本中,而是来自执行边界状态——即模型即将生成或延续目标行为的时刻。这一前后实现区分解释了为何基于答案的源有时有效:其有用成分对齐于执行边界方向,而非仅依赖目标出现。基于此,提出尾部减法(tail subtraction),从边界状态中移除共享的提示与续写语义,获得更干净、更稳定的引导信号。结果表明,引导依赖于模型即将执行的表征,而非已有内容。

原文摘要 · Abstract (English)

Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.

模型干预激活引导执行边界信号优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。