揭示大模型推理优化中一种独特训练轨迹的几何特性
On the Geometry of On-Policy Distillation

- 通过参数空间分析发现OPD更新集中在低维子空间
- 其更新方向既不像微调那样集中,也不如强化学习那般受约束
- 适合研究高效模型优化与训练机制设计的学者
在大语言模型推理优化中,基于策略的蒸馏(OPD)正日益流行,但其训练动态仍不明确。本文通过参数空间诊断,刻画了OPD更新轨迹,并与监督微调(SFT)和可验证奖励强化学习(RLVR)进行对比。结果表明,OPD处于一种松弛的非主方向状态:相比SFT,其更新影响更少权重且避开主方向更明显;相比RLVR,更新约束更宽松。此外,OPD表现出子空间锁定现象:累积更新迅速进入狭窄的低维通道。将训练限制在早期形成的更新子空间内,可保持OPD性能,但会显著降低SFT表现,说明该锁定子空间对OPD已足够。控制实验显示,稀疏化更新标记或离策略生成回溯不影响秩动态,而混合OPD与RLVR目标则会改变动态。整体表明,OPD并非SFT与RLVR之间的过渡状态,而是形成了自身独特的参数空间更新几何。
原文摘要 · Abstract (English)
On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of OPD updates in parameter space and compare it with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). A suite of parameter-space diagnostics consistently places OPD in a relaxed off-principal regime: compared with SFT, its updates affect fewer weights and avoid principal directions more strongly, while compared with RLVR, they remain less tightly constrained. Beyond this static localization, OPD exhibits subspace locking: its cumulative updates rapidly enter a narrow low-dimensional channel. Constraining training to the update subspace formed early in training preserves OPD performance but substantially degrades SFT, indicating that the locked subspace is functionally sufficient for OPD. Control experiments further show that sparsifying the update tokens and shifting rollout generation off-policy preserve the rank dynamics, whereas mixing the OPD objective with RLVR changes them. Overall, these results suggest that OPD is not merely an intermediate point between SFT and RLVR, but induces its own update geometry in parameter space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。