arXiv:2509.06608cs.LG2025-09被引 8

用轻量向量模拟推理训练,揭示模型内部计算变化机制

Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors

论文配图:Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
图 1 · 摘自论文原文
  • 通过强化学习训练插入残差流的小型控制向量
  • 最后层向量提升首词生成概率,如"To"和"Step"
  • 向量可迁移至同系列其他模型,适合研究推理机制

推理训练如何重塑大语言模型的内部计算机制仍不明确。本文研究在基础模型残差流中插入的轻量级控制向量,并通过强化学习目标进行训练。这些向量解释了大部分全微调带来的性能提升,同时保持小规模、可解释的添加性干预特性。我们发现:(i) 最后一层的控制向量表现为一种集中在首个生成标记上的令牌替换偏差,持续提升"To"和"Step"等词的概率;(ii) 前一层向量基本保持注意力模式不变,而是通过MLP和反嵌入模块起作用,优先增强过程词与结构符号;(iii) 控制向量可在同一模型家族间实现迁移。上述结果深化了对训练后控制向量如何塑造计算的理解,为激活工程及推理模型研究提供指导。

原文摘要 · Abstract (English)

The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These vectors explain a large portion of full fine-tuning performance increase while preserving the interpretability of small, additive interventions. We find that (i) the last-layer steering vector acts like a token-substitution bias concentrated on the first generated token, consistently boosting tokens such as "To" and "Step"; (ii) the penultimate-layer vector leaves attention patterns largely intact and instead operates through the MLP and unembedding, preferentially up-weighting process words and structure symbols; and (iii) the steering vectors transfer to other models from the same family. Taken together, these results deepen understanding of how trained steering vectors shape computation and should inform future work in activation engineering and the study of reasoning models.

控制向量推理机制激活工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。