揭示分组强化学习隐藏的梯度偏差,为大模型训练提供理论改进方向
On the Hidden Objective Biases of Group-based Reinforcement Learning
- 通过统一代理框架分析分组强化学习的内在机制
- 发现非均匀分组加权导致共享前缀梯度偏移
- 适用于关注大语言模型训练稳定性的研究者
分组强化学习方法(如GRPO)被广泛用于大语言模型的后训练。尽管表现优异,但其奖励优化与训练目标存在结构性不匹配。本文通过统一代理框架对GRPO类方法进行理论分析,揭示三类共性问题:(i) 非均匀分组加权会引发共享前缀词元的系统性梯度偏差;(ii) 与AdamW优化器的交互使训练动态对奖励缩放不敏感;(iii) 优化器动量在重复优化步骤下可能使策略更新超出预期截断区域。这些发现揭示了当前方法的根本局限,并为未来设计提供原则性指导。
原文摘要 · Abstract (English)
Group-based reinforcement learning methods, like Group Relative Policy Optimization (GRPO), are widely used nowadays to post-train large language models. Despite their empirical success, they exhibit structural mismatches between reward optimization and the underlying training objective. In this paper, we present a theoretical analysis of GRPO style methods by studying them within a unified surrogate formulation. This perspective reveals recurring properties that affect all the methods under analysis: (i) non-uniform group weighting induces systematic gradient biases on shared prefix tokens; (ii) interactions with the AdamW optimizer make training dynamics largely insensitive to reward scaling; and (iii) optimizer momentum can push policy updates beyond the intended clipping region under repeated optimization steps. We believe that these findings highlight fundamental limitations of current approaches and provide principled guidance for the design of future formulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。