arXiv:2507.12638cs.LG2025-07被引 26

推理微调让基础模型重用已有表征,激活新推理能力。

Reasoning-Finetuning Repurposes Latent Representations in Base Models

  • 发现基础模型中存在可被微调模型复用的潜在方向
  • 该方向在基座模型中不引发回溯,在推理模型中却能诱发
  • 为理解推理能力涌现提供了表征重用的新视角

推理微调引发的回溯行为是推理模型能力提升的关键机制。本文发现,DeepSeek-R1-Distill-Llama-8B 中回溯现象的部分原因,源于基座模型激活中已存在的一个可被重新利用的方向。具体而言,我们识别出 Llama-3.1-8B 残差流中的一个方向,当用于引导蒸馏后的推理模型时,会系统性诱发回溯行为,且其效果无法由词元层面属性简单解释。进一步发现,该方向在基座模型中并不引发回溯,表明推理微调过程对已有表征进行了再利用,构建了新的行为通路。我们还推测,该方向可能是多个协同作用方向之一,共同介导回溯行为。研究结果支持推理微调模型通过重用而非从零学习来获得新能力的观点。

原文摘要 · Abstract (English)

Backtracking, an emergent behavior elicited by reasoning fine-tuning, has been shown to be a key mechanism in reasoning models' enhanced capabilities. Prior work has succeeded in manipulating this behavior via steering vectors, but the underlying mechanism remains poorly understood. In this work, we show that the emergence of backtracking in DeepSeek-R1-Distill-Llama-8B is in part driven by a repurposed direction already present in base model activations. Specifically, we identify a direction in base Llama-3.1-8B's residual stream which systematically induces backtracking when used to steer the distilled reasoning model, and find that the effects of steering with this direction cannot be trivially explained by token-level attributes. We further find that this direction does not induce backtracking in the base model, suggesting that the reasoning finetuning process repurposes pre-existing representations to form new behavioral circuits. Additionally, we hypothesize that this direction is one of several which may work together to mediate backtracking. Our findings offer a compelling picture that reasoning-finetuned models repurpose pre-existing base model representations, rather than learn new capabilities from scratch.

推理能力表征重用微调机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。