OPD提升推理效率,但未必真正扩展模型能力边界。
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

- 通过测试时采样预算变化分析OPD效果
- 高采样数下基础模型反而更优,显示能力未真正扩展
- 适合关注推理效率而非能力边界的实践者
在多个OPD变体中,我们发现经过OPD训练的模型在不同采样预算K下保持更高的avg@K表现,而pass@K的优势随K增大逐渐转移到预训练基线模型。这表明OPD主要提升采样效率,而非持续扩展学生模型的推理能力边界。进一步分析显示,训练过程导致小K性能增强,但大K能力下降。以pass@1024为标准的问题可解性分析揭示不对称性:更多原本可解的问题变为不可解,而可解问题数量增加有限。因此,从能力扩展角度看,OPD更像一种‘幻觉式蒸馏’,其表面收益主要来自采样效率提升,而非真正获取教师模型的新能力。
原文摘要 · Abstract (English)
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding capabilities beyond the pre-OPD base model. In this study, we examine this view through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K and avg@K. Specifically, across several OPD variants, we observe that OPD-trained models maintain superior avg@K performance across sampling budgets, while the advantage in pass@K gradually shifts to the pre-OPD base models as K increases. These results suggest that OPD primarily improves sampling efficiency rather than consistently expanding the student's reasoning capability boundary. The pass@K dynamics throughout OPD training further reveal a progressive shift toward stronger small-K performance at the expense of the large-K capability boundary. Furthermore, a problem-level solvability analysis using pass@1024 as the criterion reveals an asymmetry: OPD causes more previously solvable problems to become unsolvable than previously unsolvable problems to become solvable. Together, these findings suggest that, from the perspective of capability expansion, OPD behaves more like an "illusory distillation": its apparent gains arise primarily from improved sampling efficiency rather than from acquiring genuinely new reasoning capabilities from the teacher.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。