arXiv:2603.01162cs.LGstat.ML2026-03被引 16

揭示GRPO的梯度本质是统计学中的U-统计量,解释其为何高效且可优化。

Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic

  • 从经典U-统计量视角统一理解GRPO,揭示其梯度结构
  • 证明GRPO渐近等价于理想梯度算法,性能逼近最优
  • 给出通用缩放定律,指导最优分组大小选择

Group Relative Policy Optimization(GRPO)作为DeepSeekMath和DeepSeek-R1的核心方法,已成为提升大语言模型推理能力的关键。尽管广泛应用并催生大量后续工作,其理论性质仍不清晰。本文通过经典U-统计量框架,揭示GRPO的策略梯度本质上是U-统计量,可刻画其均方误差(MSE),推导出有限样本误差界与次优性差距的渐近分布。研究发现,GRPO渐近等价于一个拥有理想价值函数的基准算法——该函数可在每轮训练中准确评估当前策略质量,并在一大类策略梯度算法中实现渐近最优性能。此外,我们建立了通用缩放定律,为最优分组大小的选择提供理论指导。实验验证了理论结论,表明最优分组大小具有普适性,并证实了GRPO的‘理想’属性。

原文摘要 · Abstract (English)

Group relative policy optimization (GRPO), a core methodological component of DeepSeekMath and DeepSeek-R1, has emerged as a cornerstone for scaling reasoning capabilities of large language models. Despite its widespread adoption and the proliferation of follow-up works, the theoretical properties of GRPO remain less studied. This paper provides a unified framework to understand GRPO through the lens of classical U-statistics. We demonstrate that the GRPO policy gradient is inherently a U-statistic, allowing us to characterize its mean squared error (MSE), derive the finite-sample error bound and asymptotic distribution of the suboptimality gap for its learned policy. Our findings reveal that GRPO is asymptotically equivalent to an oracle policy gradient algorithm -- one with access to a value function that quantifies the goodness of its learning policy at each training iteration -- and achieves asymptotically optimal performance within a broad class of policy gradient algorithms. Furthermore, we establish a universal scaling law that offers principled guidance for selecting the optimal group size. Empirical experiments further validate our theoretical findings, demonstrating that the optimal group size is universal, and verify the oracle property of GRPO.

强化学习策略梯度理论分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。