揭示大模型在线蒸馏成功的关键条件与机制,提出有效修复方案。
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

- 发现蒸馏成败取决于师生思维模式匹配及教师提供新能力
- 成功蒸馏时高概率词元逐步对齐,共享词元占97%-99%概率质量
- 提出冷启动与提示对齐策略可挽救失败的蒸馏,适合模型优化者
在线蒸馏(OPD)已成为大语言模型后训练的核心技术,但其训练动态仍不清晰。本文系统研究了OPD的动态与机制,首次识别出两项决定成败的关键条件:(i) 学生与教师需具备兼容的思维模式;(ii) 即使思维一致且得分更高,教师也必须提供学生训练中未见的新能力。通过弱到强反向蒸馏验证,发现同家族1.5B与7B教师在分布上对学生不可区分。深入分析词元级机制发现,成功蒸馏表现为在学生访问状态下的高概率词元逐步对齐,且共享词元集集中了97%-99%的概率质量。进一步提出两种实用策略恢复失败蒸馏:离线冷启动和教师对齐提示选择。最后指出,OPD看似免费的密集词元奖励实则有代价,引发其能否扩展至长序列蒸馏的疑问。
原文摘要 · Abstract (English)
On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly understood. This paper provides a systematic investigation of OPD dynamics and mechanisms. We first identify that two conditions govern whether OPD succeeds or fails: (i) the student and teacher should share compatible thinking patterns; and (ii) even with consistent thinking patterns and higher scores, the teacher must offer genuinely new capabilities beyond what the student has seen during training. We validate these findings through weak-to-strong reverse distillation, showing that same-family 1.5B and 7B teachers are distributionally indistinguishable from the student's perspective. Probing into the token-level mechanism, we show that successful OPD is characterized by progressive alignment on high-probability tokens at student-visited states, a small shared token set that concentrates most of the probability mass (97%-99%). We further propose two practical strategies to recover failing OPD: off-policy cold start and teacher-aligned prompt selection. Finally, we show that OPD's apparent free lunch of dense token-level reward comes at a cost, raising the question of whether OPD can scale to long-horizon distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。