arXiv:2605.05892cs.CLcs.LG2026-05

提出新型激活调控方法,让模型推理时更精准可控。

Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention

论文配图:Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
图 1 · 摘自论文原文
  • 用可学习的流场代替固定变换,实现动态、多步调节
  • 在AxBench上超越提示词,最高得分达1.113(无概念微调)
  • 适合需要高精度推理干预的研究者与应用开发

激活调控作为一种在推理阶段控制语言模型行为的新兴方法,通过修改中间表示而保持模型参数不变。然而大规模评估如AxBench显示,现有调控方法常被简单提示词超越,且对未见概念泛化能力差。我们假设这些局限源于先前方法共有的未经验证简化假设——通常将调控限制为固定、单步、位置无关的变换。本文提出FLAS(基于流的激活调控),学习一个通用的、由概念条件决定的速度场 $v_t(h,t,c)$,无需依赖上述假设即可将原始激活映射至目标激活。在AxBench上,FLAS是首个持续优于提示词的可学习方法,在Gemma-2-2B-IT和Gemma-2-9B-IT上分别达到1.015和1.113的保留均值,且无需针对每个概念进行微调。对学习到的流场分析显示其轨迹呈曲线、多步、随标记变化,暗示此前关于激活空间几何的认知可能不完整。

原文摘要 · Abstract (English)

Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model parameters frozen. However, large-scale evaluations such as AxBench show that existing steering methods are often outperformed by simple in-context prompting and generalize poorly to unseen concepts. We hypothesize that these limitations arise from unvalidated simplifying assumptions shared across prior methods, which typically restrict steering interventions to fixed, single-step, position-invariant transforms. We propose FLAS (Flow-based Activation Steering), which learns a general, concept-conditioned velocity field $v_t(h,t,c)$ that transports unsteered activations to steered ones without relying on these assumptions. On AxBench, FLAS is the first learned method to consistently outperform prompting, reaching held-out harmonic means of $1.015$ on Gemma-2-2B-IT and $1.113$ on Gemma-2-9B-IT without per-concept tuning. Analysis of the learned flow shows curved, multi-step, token-varying trajectories, which suggests that previous hypotheses on activation space geometry might be incomplete.

推理干预激活调控流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。