arXiv:2604.14090cs.CL2026-04ACL被引 4

将推理时的激活值调控视为一种新型模型适应方式。

From Weights to Activations: Is Steering the Next Frontier of Adaptation?

  • 提出以激活空间干预为核心的新适应范式
  • 证明该方法可局部、可逆地改变模型行为
  • 适合关注模型可解释性与轻量级调优的研究者

语言模型的训练后适应通常通过参数更新或基于输入的方法(如微调、高效参数适配和提示)实现。与此同时,越来越多的工作在推理时修改内部激活值以影响模型行为,这种方法称为‘引导’(steering)。尽管应用日益广泛,但引导方法很少被置于与传统适应方法相同的分析框架中。本文主张将引导视为一种模型适应形式,引入一套功能标准来评估适应方法,并以此比较引导与经典方法。分析表明,引导是一种基于激活空间的靶向干预的独立适应范式,可在不更新参数的情况下实现局部且可逆的行为调整。这一框架厘清了引导与现有方法的关系,推动建立统一的模型适应分类体系。

原文摘要 · Abstract (English)

Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficient adaptation, and prompting. In parallel, a growing body of work modifies internal activations at inference time to influence model behavior, an approach known as steering. Despite increasing use, steering is rarely analyzed within the same conceptual framework as established adaptation methods. In this work, we argue that steering should be regarded as a form of model adaptation. We introduce a set of functional criteria for adaptation methods and use them to compare steering approaches with classical alternatives. This analysis positions steering as a distinct adaptation paradigm based on targeted interventions in activation space, enabling local and reversible behavioral change without parameter updates. The resulting framing clarifies how steering relates to existing methods, motivating a unified taxonomy for model adaptation.

模型适应激活调控推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。