剖析自蒸馏中推理路径坍缩的三个调控因素,厘清方法本质与瓶颈。
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
- 用模型自身生成的带偏见信号进行自蒸馏,通过三类调控机制影响性能。
- 在数学推理任务中,自蒸馏可达到强化学习水平,但易引发推理路径坍缩。
- 适合关注大模型自我训练机制与失败模式的研究者阅读。
在线策略自蒸馏(OPSD)让语言模型基于自身生成结果进行训练,无需额外教师模型。教师是模型自身,仅依赖学生在测试时无法获得的特权信息(如参考解、计划或环境反馈),虽不更强但更知情。早期结果表明其准确率接近强化学习,且生成样本量仅为后者的几分之一。然而,这种信息不对称导致信号偏差,引发主流问题:坍缩——模型逐步缩小可产生的推理路径范围。该问题并非仅限于OPSD,但特权信息会加剧其发生。本文将坍缩视为由三个杠杆控制的症状:(i) 信号作用位置(令牌权重分配方式);(ii) 教师所见内容(特权信息类型);(iii) 信号变化时机(教师动态与引导衰减)。研究聚焦数学推理领域,不报告新实验,贡献在于建立统一术语体系,区分已确定结论与仍存争议的问题。
原文摘要 · Abstract (English)
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。