发现自蒸馏性能可提前预测,无需训练即可优化配置。
A Predictive Law for On-Policy Self-Distillation From World Feedback
- 通过初始性能差距线性预测最终自蒸馏效果。
- 该规律在不同模型和场景下均成立,跨规模有效。
- 适合想高效利用世界反馈的强化学习研究者。
从简单标量奖励转向更丰富的世界反馈,是实现可扩展强化学习后训练的自然路径。最近提出的基于策略自蒸馏(OPSD)方法利用任意反馈作为学习信号,但其可靠性与GRPO等成熟方法相比尚不明确。我们发现初始学生-教师性能差距与最终性能提升之间存在显著且一致的线性相关性,这一关系在不同上下文类型和模型家族中均成立,为预测未训练的OPSD配置结果提供了强大工具。有趣的是,该线性可预测性在模型规模上保持稳定,暗示了未来在更大模型、更强上下文学习能力基础上建立新经验缩放定律的可能性。本质上,我们的发现表明可在训练前预测并调优OPSD性能,为将世界反馈作为后训练流程核心组件提供原则性方法。
原文摘要 · Abstract (English)
Moving beyond simple scalar rewards toward richer world feedback is a natural path to more scalable RL post-training. On-policy self-distillation (OPSD) is a promising recent approach that uses arbitrary feedback as learning signal, yet its reliability compared to established methods, such as GRPO, remains unclear. We identify a strikingly consistent linear correlation between the initial student-self-teacher performance gap and the final performance improvement in OPSD. This relationship holds across context types and model families, providing a powerful predictive law for anticipating the outcome of an OPSD configuration without running the full training procedure. Interestingly, we show that this linear predictability holds with model scale, suggesting a potential basis for new empirical scaling laws on larger models with stronger in-context learning capabilities. In essence, our findings show that OPSD performance can be predicted and tuned before training, offering a principled way to incorporate world feedback as a first-class component of the post-training pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。