arXiv:2605.06188cs.AIcs.CL2026-05被引 3

OPSD更适合压缩推理模型输出而非纠错,提出三阶段训练新流程

OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models

  • 分离正确与错误推理路径,分别应用OPSD验证其作用机制
  • 仅在正确路径上训练可显著缩短响应且不降低准确率
  • 建议采用SFT→RLVR→OPSD的后训练流程,提升数学推理效率

最近提出的基于策略自蒸馏(OPSD)作为强化学习结合可验证奖励(RLVR)的替代方案,通过利用特权上下文条件下的自教师,在词元级别实现信用分配,有望提高准确率并缩短输出。然而,在启用思考的数学推理任务中,其准确率提升效果减弱甚至转为下降。我们假设:事后监督在短推理输出中能提供更好的词元级替代,但在长推理轨迹中更易识别冗余而非提供更优替换。为此,我们将OPSD分别应用于正确与错误推理轨迹组,以孤立观察压缩与修正效果。结果表明,在启用思考的数学推理中,OPSD主要表现为压缩机制而非修正机制:仅在正确轨迹上训练可有效缩短输出且保持准确率;而在错误轨迹上训练则会损害准确率。基于此,我们提出改进的后训练流程:SFT → RLVR → OPSD。

原文摘要 · Abstract (English)

On-Policy Self-Distillation (OPSD) has recently emerged as an alternative to Reinforcement Learning with Verifiable Rewards (RLVR), promising higher accuracy and shorter responses through token-level credit assignment from a self-teacher conditioned on privileged context. However, this promise does not carry over to thinking-enabled mathematical reasoning, where reported accuracy gains shrink and sometimes turn negative. We hypothesize that hindsight supervision can specify better token-level alternatives in short thinking-disabled outputs, but in long thinking-enabled traces it more readily identifies redundancy than supplies better replacements. To test this, we applied OPSD separately to correct and incorrect rollout groups, so that compression and correction can be observed in isolation. Our results show that in thinking-enabled mathematical reasoning, OPSD behaves most reliably as a compression mechanism rather than a correction mechanism: training only on correct rollouts preserves accuracy while substantially shortening responses, whereas training only on incorrect rollouts damages accuracy. In light of these findings, we propose a revised post-training pipeline for thinking-enabled mathematical reasoning: SFT then RLVR then OPSD.

推理模型自蒸馏强化学习数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。