arXiv:2605.17231cs.LGcs.CL2026-05被引 1

用几何修正方法让模型改写语法更精准,不扰其他功能。

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

论文配图:FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
图 1 · 摘自论文原文
  • 基于费雪信息度量推导出最优激活调控方向
  • 在多个语法概念上使副作用降低1.4至6.5倍
  • 适合需要精细控制语言模型行为的研究者

激活调控已成为无需更新参数即可修改语言模型行为的轻量级方法,但现有方法在不同层间不稳定,易干扰非目标概念。我们发现主流方法(如CAA、ActAdd、ITI)隐含假设中间激活空间为欧氏空间,这一假设根本错误。实际决定隐藏状态扰动如何影响输出的是软最大层的费雪信息度量,通过中间层雅可比矩阵拉回至中间层。据此推导出闭式调控方向,在中间层施加扰动可最小化非目标偏差。该框架在早期和中间层最有效,此时度量明显非欧氏。我们在三个动词形态概念(第三人称单数、进行时、过去时)上评估,使用标准反事实概念评测。在GPT-2 Small上,该方法使非目标KL散度中位下降1.4–6.5倍;在Llama-3-8B和Qwen3-8B上,早期与中间层中位下降1.8–3.6倍。结果表明,即使在更大更复杂的模型上,几何修正仍具优势。

原文摘要 · Abstract (English)

Activation steering has emerged as a lightweight approach for modifying language model behavior without parameter updates, yet existing methods remain brittle: unstable across layers and prone to disturbing behavior unrelated to the target concept. We trace these failures to a hidden assumption shared by widely-used methods such as CAA, ActAdd, and ITI: that the intermediate activation space is Euclidean. We show this assumption is fundamentally flawed. The metric that actually governs how a hidden-state perturbation changes the output is the Fisher information metric of the softmax layer, pulled back to the intermediate layer through the Jacobian of the intervening layers. From it we derive a closed-form steering direction, applied to a hidden state at an intermediate layer, that reaches a target concept change with the least non-target distortion. The framework is sharpest in the early and middle intermediate layers, where the metric is strongly non-Euclidean and geometric correction matters most. We evaluate it on three verb-morphology concepts: third-person-singular, progressive, and past-tense inflection, following standard counterfactual-concept evaluation. On GPT-2 Small, this non-Euclidean geometry is borne out empirically, and our method lowers off-target KL divergence by median factors of 1.4--6.5x against individual steering baselines. On Llama-3-8B and Qwen3-8B, it lowers off-target KL by median factors of 1.8--3.6x against individual baselines at the early and middle layers. These results show that geometric correction retains its advantage on larger models with more complex internal structure.

激活调控几何修正语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。