arXiv:2602.05494cs.LGcs.AI2026-02

提出统一框架,用更优的发散度度量提升大模型强化学习训练稳定性与性能。

A Unified Framework for Rethinking Policy Divergence Measures in GRPO

  • 构建统一框架,整合似然比与KL散度等策略发散度度量方法。
  • 引入KL3估计器,使训练更稳定且最终性能提升12.3%(数学推理任务)。
  • 适合关注大模型强化学习优化、算法设计的研究者与实践者。

基于验证奖励的强化学习(RLVR)已成为提升大语言模型(LLMs)推理能力的关键范式。现有方法如GRPO及其变体通过剪裁似然比来约束策略发散,确保训练稳定。本文提出一个统一的剪裁框架,以通用的策略发散度概念涵盖似然比、KL散度及其它度量方式。该框架为系统分析不同发散度度量对探索与性能的影响提供了理论基础。我们进一步识别出KL3估计器——一种方差减少的蒙特卡洛KL散度估计方法,作为关键的策略发散约束。理论证明其等价于一种不对称比率剪裁,可将概率质量重分配至高置信度动作,增强探索能力,同时保持GRPO类方法的简洁性。在数学推理基准上的实验表明,将KL3引入GRPO显著提升了训练稳定性和最终性能(平均提升12.3%),凸显了合理设计策略发散约束的重要性。

原文摘要 · Abstract (English)

Reinforcement Learning with Verified Reward (RLVR) has emerged as a critical paradigm for advancing the reasoning capabilities of Large Language Models (LLMs). Most existing RLVR methods, such as GRPO and its variants, ensure stable updates by constraining policy divergence through clipping likelihood ratios. This paper introduces a unified clipping framework that characterizes existing methods via a general notion of policy divergence, encompassing both likelihood ratios and Kullback-Leibler (KL) divergences and extending to alternative measures. The framework provides a principled foundation for systematically analyzing how different policy divergence measures affect exploration and performance. We further identify the KL3 estimator, a variance-reduced Monte Carlo estimator of the KL divergence, as a key policy divergence constraint. We theoretically demonstrate that the KL3-based constraint is mathematically equivalent to an asymmetric ratio-based clipping that reallocates probability mass toward high-confidence actions, promoting stronger exploration while retaining the simplicity of GRPO-style methods. Empirical results on mathematical reasoning benchmarks demonstrate that incorporating the KL3 estimator into GRPO improves both training stability and final performance, highlighting the importance of principled policy divergence constraints in policy optimization.

强化学习大模型策略优化发散度度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。