arXiv:2605.10889cs.LGcs.AI2026-05被引 5

提出可逐令牌诊断强化学习蒸馏效果的新方法。

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

  • 设计逐令牌、逐问题的无训练诊断框架,评估蒸馏信号质量。
  • 发现错误推理路径上蒸馏信号比正确路径更接近理想梯度。
  • 揭示蒸馏最佳上下文依赖学生模型能力与任务类型,无通用最优解。

在策略蒸馏为推理模型训练提供了密集的逐标记监督信号;然而,这种信号在何种条件下有益或有害仍不明确。应使用何种教师模型?自蒸馏时,哪个具体上下文应作为监督信号?最优选择是否随标记变化?目前解决这些问题通常需要代价高昂的训练实验,其整体性能指标掩盖了单个标记层面的动态。本文提出一种无需训练的诊断框架,可在最高分辨率下(逐标记、逐问题、逐教师)运行。我们推导出理想的每节点梯度——即最大化学生成功概率的参数更新。进而开发了一种可扩展的目标展开算法,高效估计该梯度,即使在长中间思考链中亦可实现。梯度对齐得分(理想梯度与给定蒸馏梯度之间的余弦相似度)量化了特定配置与理想信号的接近程度。在多种自蒸馏设置和外部教师模型下,我们发现蒸馏指导信号在错误展开路径上的对齐度显著高于正确路径,在后者中学生已表现良好,教师信号趋于噪声化。此外,我们发现最优蒸馏上下文同时依赖学生模型容量与目标任务,且不存在单一普适有效的配置。这些发现推动了针对任务和标记的逐级诊断分析在蒸馏中的应用。

原文摘要 · Abstract (English)

On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically requires costly training runs whose aggregate performance metrics obscure the dynamics at the level of individual tokens. We introduce a training-free diagnostic framework that operates at the highest resolution: per token, per question, and per teacher. We derive an ideal per-node gradient defined as the parameter update that maximally increases the student's probability of success. We then develop a scalable targeted-rollout algorithm to estimate this gradient efficiently, even for long chains of intermediate thoughts. The gradient alignment score, defined as the cosine similarity between this ideal gradient and any given distillation gradient, quantifies the extent to which a particular configuration approximates the ideal signal. Across a range of self-distillation settings and external teacher models, we observe that distillation guidance exhibits substantially higher alignment with the ideal on incorrect rollouts than on correct ones, where the student already performs well and the teacher's signal tends to become noisy. Furthermore, we find that the optimal distillation context depends jointly on the student model's capacity and the target task, and that no single universally effective configuration emerges. These findings motivate the use of per-task, per-token diagnostic analyses for distillation.

强化学习模型蒸馏诊断分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。