arXiv:2606.00305cs.CLcs.AI2026-06被引 4

通过未来轨迹信息提升学生模型推理一致性

Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance

论文配图:Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
图 1 · 摘自论文原文
  • 利用近未来轨迹判断真实推理分歧点,避免误修表面错误
  • 在AIME24上准确率从60.0%提升至63.3%,总体达52.2%
  • 适合需要稳定推理链的复杂任务场景

On-Policy Distillation (OPD) 通过教师监督下采样自身体策略的轨迹来改进大语言模型的推理能力。尽管基于轨迹采样,其学习信号仍为逐标记级别:通过高损失标记识别偏差,并用局部反KL修正。我们发现这种‘轨迹采样但标记学习’机制难以可靠地将学生轨迹引导至教师轨迹。约30%的高损失标记处于低差异区间,表明许多仅为表面形式不匹配而非真正推理分叉。此外,即使真正发散的标记也难通过孤立标记监督修复,因推理失败常表现为短时程分布漂移。我们提出轨迹感知式OPD(TOPD),利用近未来轨迹信息识别真实发散状态,并将指导分布到多个未来标记。实验显示,抑制非发散高损失标记使标准OPD准确率从47.8%提升至48.2%,而TOPD进一步提升至52.2%,在AIME24上从60.0%升至63.3%,AIME25上从46.7%升至53.3%。

原文摘要 · Abstract (English)

On-Policy Distillation (OPD) improves large language model reasoning by training a student model on trajectories sampled from its own policy under teacher supervision. Although OPD operates on trajectories, its learning signal remains token-level: it identifies deviations through high-loss tokens and repairs them through local reverse-KL correction. We show that this "trajectory-sampled but token-learned" mechanism cannot reliably bridge student trajectories toward teacher trajectories. About 30% of high-loss tokens fall into the low-divergence regime, indicating that many are surface-form mismatches rather than real reasoning forks. Moreover, even truly divergent tokens are difficult to repair with isolated token-level supervision, since reasoning failures often unfold as short-horizon distributional drift. We propose Trajectory-aware OPD (TOPD), which uses near-future trajectory information to identify real divergent states and distribute guidance across multiple future tokens. Experiments show that suppressing non-divergent high-loss tokens improves standard OPD from 47.8% to 48.2% average accuracy, while TOPD further improves performance to 52.2%, with gains on AIME24 from 60.0% to 63.3% and AIME25 from 46.7% to 53.3%.

推理增强轨迹引导知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。