arXiv:2606.09304cs.CLcs.LG2026-06被引 8

通过符号一致性门控提升学生模型在自洽强化学习中的表现

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

论文配图:SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
图 1 · 摘自论文原文
  • 用二元验证器作为教师可信度信号,分阶段引入可靠轨迹
  • 在符号一致处直接更新,不一致处插值,提升训练稳定性
  • 数学推理任务上平均提升1.98(样本级)和7.50(题目级)

在线策略蒸馏(OPD)利用更强的教师模型对学生的自身轨迹进行密集的逐标记监督,通常优于离线蒸馏和标准强化学习。然而我们发现其有效性依赖于两个常被打破的假设:学生与教师轨迹层面的对齐,以及教师偏好在标记层面的一致可靠性。为此,我们提出符号门控在线策略蒸馏(SG-OPD),通过二元验证器作为信任信号,在两个互补粒度上发挥作用:冷启动阶段分阶段引入验证器认可的教师轨迹;符号一致性门控则在教师与验证器方向一致时直接外推更新,在不一致时进行插值。在竞赛级数学推理基准上的实验表明,SG-OPD持续优于标准OPD,样本级平均提升1.98,题目级平均提升7.50。

原文摘要 · Abstract (English)

On-policy distillation (OPD) trains a student on its own trajectories with dense per-token supervision from a stronger teacher, and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its effectiveness implicitly relies on two assumptions that frequently break in practice: trajectory-level alignment between the student and the teacher, and uniform token-level reliability of the teacher's preferences. We therefore propose Sign-Gated On-Policy Distillation (SG-OPD), which uses a binary verifier as a trust signal for the teacher at two complementary granularities: phased teacher sampling mixes in verifier-endorsed teacher rollouts at cold-start, and a sign-consistency gate extrapolates the distillation update on tokens where the teacher agrees with the verifier-correct direction and interpolates it where it disagrees. Experiments on competition-level mathematical reasoning benchmarks show that SG-OPD consistently outperforms standard OPD, with average gains of 1.98 and 7.50 at the per-sample and per-question levels, respectively.

强化学习知识蒸馏符号一致性数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。