提出新型强化学习蒸馏方法,提升模型在复杂任务中的稳定性和适应性。
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

- 用局部差异最小化替代无效负反馈,优化策略更新机制。
- 在数学推理任务中,初始化强弱时分别提升4.09和8.34分。
- 适合需要持续学习与工具调用的复杂决策场景。
基于策略的蒸馏(OPD)通过教师在词元级提供反馈训练学生模型,通常优于离策略蒸馏和标准强化学习。然而我们发现其标准优势加权策略梯度存在三大结构性缺陷:更新方差高、零优势区域梯度消失,以及纠正信号不足时探索受阻。为此提出非对称在线策略蒸馏(AOPD),在非正优势区域以局部差异最小化替代无效负强化,同时保留正向强化学习。在数学推理基准上的实验表明,AOPD在强/弱初始化下平均性能分别提升4.09 / 8.34分。此外,AOPD在训练过程中保持更高策略熵,并在连续工具使用适应中展现更强能力保留性。
原文摘要 · Abstract (English)
On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its standard advantage weighted policy gradient suffers from three structural weaknesses, including high variance updates, vanishing gradients in zero-advantage regions, and exploration bottlenecks when corrective signals are insufficient. We therefore propose Asymmetric On-Policy Distillation (AOPD), which replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning. Experiments on mathematical reasoning benchmarks show that AOPD consistently outperforms standard OPD, with average gains of 4.09 / 8.34 under strong / weak initialization, respectively. AOPD also maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。