提出激活调控与权重更新等价框架,实现高效模型微调
Weight Updates as Activation Shifts: A Principled Framework for Steering
- 发现激活空间干预与权重更新在数学上等价,可指导调控位置选择
- 后块输出处调控仅用0.04%参数即达全量微调99%以上精度
- 联合双空间优化突破单一方法性能上限,适合资源受限场景
激活调控有望成为极高效的模型适配方式,但其效果依赖于干预位置和参数化等关键设计,而这些目前仍依赖经验启发而非理论基础。本文首次建立激活空间干预与权重空间更新的一阶等价关系,推导出激活调控复现微调行为的条件。该等价性为调控设计提供了理论依据,并识别出后块输出为理论上最优且高度表达力的干预点。进一步解释了为何某些位置优于其他位置,并表明权重更新与激活更新具有不同且互补的功能角色。基于此分析,提出联合适应新方法,在两个空间同步训练。后块调控在多数任务和模型上平均仅需0.04%参数即可达到全参数微调0.2%-0.9%误差范围内,显著优于ReFT、PEFT及LoRA等现有方法。最终证明联合适应常超越单独使用权重或激活更新的性能上限,开启高效模型适配的新范式。
原文摘要 · Abstract (English)
Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices -- such as intervention location and parameterization -- that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions and weight-space updates, deriving the conditions under which activation steering can replicate fine-tuning behavior. This equivalence yields a principled framework for steering design and identifies the post-block output as a theoretically-backed and highly expressive intervention site. We further explain why certain intervention locations outperform others and show that weight updates and activation updates play distinct, complementary functional roles. This analysis motivates a new approach -- joint adaptation -- that trains in both spaces simultaneously. Our post-block steering achieves accuracy within 0.2%-0.9%$ of full-parameter tuning, on average across tasks and models, while training only 0.04% of model parameters. It consistently outperforms prior activation steering methods such as ReFT and PEFT approaches including LoRA, while using significantly fewer parameters. Finally, we show that joint adaptation often surpasses the performance ceilings of weight and activation updates in isolation, introducing a new paradigm for efficient model adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。