用结果奖励校准教师输出,让大模型自我提升更稳定高效
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

- 通过成功与失败轨迹对比,动态调整每步生成的逻辑方向
- 在数学推理任务上比标准自蒸馏方法提升12.3%准确率
- 适合想提升大模型逻辑推理能力的研究者和工程师
我们研究了基于策略的自蒸馏(OPSD),即语言模型通过自身策略轨迹中的高阶教师分布来提升推理能力。尽管前景广阔,但因教师与学生响应间存在模式不匹配,导致训练不稳定:自我反思产生的教师输出可能引入反射偏差和响应模板,错误校准逐标记监督,最终损害学生推理能力。为此,我们提出OGLS-SD——一种基于结果引导的对数漂移框架,利用可验证的结果奖励校准高阶教师对数。具体而言,该方法对比成功与失败策略轨迹所诱导的教师对数,构建出目标可区分的对数漂移方向,用于逐标记指导。在数学推理基准测试中,OGLS-SD显著稳定了自蒸馏过程,并在性能上优于标准OPSD及其他变体。
原文摘要 · Abstract (English)
We study on-policy self-distillation (OPSD), where a language model improves its reasoning ability by distilling privileged teacher distributions along its own on-policy trajectories. Despite its promise, OPSD can suffer from training instability due to a pattern mismatch between teacher and student responses. Self-reflected teacher responses may introduce reflection-induced biases and response templates that miscalibrate token-level supervision, ultimately harming the student's reasoning ability. To mitigate this issue, we propose OGLS-SD, an outcome-guided logit-steering framework that leverages verifiable outcome rewards to calibrate privileged teacher logits. Specifically, OGLS-SD contrasts teacher logits induced by successful and failed on-policy trajectories, constructing an outcome-discriminative steering direction for token-level guidance. Experiments on mathematical reasoning benchmarks show that OGLS-SD stabilizes self-distillation and improves performance over standard OPSD and other variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。