OPSD效果未必来自特权答案,更多是教师行为变化所致。
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

- 用不同问题的解代替原参考解,测试特权信息作用
- 新方法在三数据集上仍优于基线,接近原始OPSD表现
- 适合研究强化学习中教师行为影响的学者
On-Policy Self-Distillation(OPSD)通常被理解为特权信息的传递:教师观察目标问题的验证解,并指导学生轨迹。然而,这种解释混淆了两种效应。参考解不仅揭示当前实例的答案,也改变了教师提供逐令牌监督的上下文。本文提出$$\mathrm{OP}^{2}\mathrm{SD}\u0024$(从其他问题中进行的OPSD),用另一例题的问题与解替换配对参考解,同时保持学生回放、教师模型和蒸馏目标不变。在三个模型和三个数学基准上,$$\mathrm{OP}^{2}\mathrm{SD}\u0024$均优于基线模型,且与OPSD性能相当。该结果表明,OPSD的提升并非必然源于对参考解的访问,教师的上下文诱导行为同样关键。
原文摘要 · Abstract (English)
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with $\mathrm{OP}^{2}\mathrm{SD}$ (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, $\mathrm{OP}^{2}\mathrm{SD}$ improves over the base model, remains competitive with OPSD. The success of $\mathrm{OP}^{2}\mathrm{SD}$ implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。