OPD通过迁移推理行为实现泛化,但多教师组合会引发能力冲突。
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
- OPD学习的是教师的推理逻辑而非具体答案
- 同源师生对可跨语言、跨领域泛化,异源则仅限训练分布
- 多教师融合时会出现能力此消彼长的权衡现象
在策略蒸馏(OPD)中,学生通过监督自身采样轨迹来迁移教师能力,但其泛化行为仍不清晰。现有研究多在单一领域和接近训练数据的基准上评估。本文通过控制变量,系统考察了从域内分布偏移、跨域迁移到多教师设置下的泛化表现。发现OPD传递的是教师的推理行为而非特定问题的答案:训练难度影响小,即使教师从未解决的问题也具价值。泛化效果强烈依赖师生来源关系——同源对能将学生推向教师在多种语言、推理深度乃至其他领域的性能水平,而异源对仅适应训练分布。这种广泛影响是双刃剑:因无法通过路由限制教师作用范围,多教师组合会导致其能力间的混合依赖性权衡。研究揭示了OPD何时泛化,并为诊断多教师设置提供了新视角。
原文摘要 · Abstract (English)
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。