arXiv:2608.18271cs.LG2026-08

研究发现,自我蒸馏中的参考信息未必提升性能,学生更多在复用模型原有推理能力。

Rethinking Privileged Information in On-Policy Self-Distillation

  • 通过分离参考信息与教师监督,分析其对学生的实际影响。
  • 学生即使无正确参考也能提升,其他问题解法甚至更优。
  • 性能提升不来自参考信息对齐,而是模型自身推理行为的恢复。

基于Qwen3模型(1.7B至8B),在科学与数学数据集上开展自蒸馏实验。研究提出分析框架,将参考信息带来的监督与无参考教师监督相分离,并测量二者对学生预测变化的对齐程度。结果显示,正确参考信息在不同生成模式、模型规模和训练数据下并未持续带来性能优势;学生可在无正确参考情况下仍实现提升,且其他问题的解决方案在多个数学推理基准上表现更佳。学生预测与基础模型的思维行为对齐度高于与参考信息的对齐度,但由其他问题构建的控制组可重现大部分对齐现象。此外,参考信息带来的更强对齐并不稳定对应性能提升。因此,性能增益与分布对齐无法可靠判断特权参考信息在自蒸馏中的实际作用。

原文摘要 · Abstract (English)

On-policy self-distillation (OPSD) trains a student on its own responses using token-level supervision from the same model conditioned on privileged reference information. We investigate whether performance gains from OPSD show that the student learned the information in the reference or instead reflect recovery of reasoning behavior already present in the base model. We perform OPSD experiments on science and mathematics datasets using Qwen3 models ranging from 1.7B to 8B. Our analysis framework separates the supervision induced by the reference from the supervision provided by the teacher without the reference and measures how each aligns with changes in the student's predictions. The correct reference does not provide a consistent performance benefit across teacher generation modes, model sizes, and training datasets. Students can improve without the correct reference, and a solution from another problem can outperform the correct solution on several mathematical reasoning benchmarks. The student's predictions align more strongly with the base model's thinking behavior than with the supervision induced by the reference, but controls constructed from other problems reproduce much of both alignments. Moreover, stronger alignment attributable to the correct reference does not reliably coincide with a greater performance benefit from the reference. Performance gains and distributional alignment alone therefore cannot determine how privileged reference information contributes to student learning in OPSD.

自蒸馏推理能力模型对齐参考信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。