arXiv:2603.11137cs.LGcs.CL2026-03被引 54

通过放松模仿约束,让小模型更高效地学习大模型的推理能力。

Scaling Reasoning Efficiently via Relaxed On-Policy Distillation

  • 用教师模型的奖励信号动态调整学生模型学习过程
  • 训练效率提升6.7~12倍,7B模型实现32B模型的视觉推理性能
  • 适合资源受限场景下需高效推理的模型部署

在线策略蒸馏在将推理能力迁移到容量受限模型中至关重要,但易出现不稳定和负迁移。我们从理论与实证上表明,在线策略蒸馏可被理解为一种策略优化,其中教师-学生对数似然比充当令牌奖励。基于此,我们提出REOPOLD(Relaxed On-Policy Distillation)框架,通过混合奖励裁剪、基于熵的令牌级动态采样及统一的探索到精炼训练策略,缓解严格模仿带来的不稳定性。实验显示,REOPOLD在数学、视觉和代理工具使用等推理任务中,相比基线方法显著提升样本效率,并增强推理时扩展能力。具体而言,其训练样本效率较近期强化学习方法高6.7~12倍,使7B学生模型在视觉推理上达到32B教师模型水平,且推理速度提升约3.32倍。

原文摘要 · Abstract (English)

On-policy distillation is pivotal for transferring reasoning capabilities to capacity-constrained models, yet remains prone to instability and negative transfer. We show that on-policy distillation can be interpreted, both theoretically and empirically, as a form of policy optimization, where the teacher-student log-likelihood ratio acts as a token reward. From this insight, we introduce REOPOLD (Relaxed On-Policy Distillation) a framework that stabilizes optimization by relaxing the strict imitation constraints of standard on-policy distillation. Specifically, REOPOLD temperately and selectively leverages rewards from the teacher through mixture-based reward clipping, entropy-based token-level dynamic sampling, and a unified exploration-to-refinement training strategy. Empirically, REOPOLD surpasses its baselines with superior sample efficiency during training and enhanced test-time scaling at inference, across mathematical, visual, and agentic tool-use reasoning tasks. Specifically, REOPOLD outperforms recent RL approaches achieving 6.7~12x greater sample efficiency and enables a 7B student to match a 32B teacher in visual reasoning with a ~3.32x inference speedup.

模型蒸馏推理优化高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。