arXiv:2505.22257cs.LGstat.ML2025-05被引 36

改进强化学习策略优化,离线版比在线版更稳定高效

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

  • 将GRPO方法拓展至离线训练,提升采样效率
  • 离线GRPO在奖励上显著优于或持平在线版本
  • 适合追求训练稳定性与资源效率的研究者

我们重新审视了组相对策略优化(GRPO)在在线与离线优化场景下的表现。受近期关于离线近端策略优化(PPO)研究的启发,该方法提升了训练稳定性、采样效率和内存使用。此外,对GRPO的分析表明,用离线样本估计优势函数可能有益。基于此,我们将GRPO适配至离线设置。结果表明,无论是在线还是离线的GRPO目标,均能提升奖励表现。这一发现促使我们在离线版本中引入截断代理目标。我们在后训练阶段采用可验证奖励评估两种变体的性能,结果显示,离线GRPO要么显著优于,要么表现持平于在线版本。

原文摘要 · Abstract (English)

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Optimization (PPO), which improves training stability, sampling efficiency, and memory usage. In addition, a recent analysis of GRPO suggests that estimating the advantage function with off-policy samples could be beneficial. Building on these observations, we adapt GRPO to the off-policy setting. We show that both on-policy and off-policy GRPO objectives yield an improvement in the reward. This result motivates the use of clipped surrogate objectives in the off-policy version of GRPO. We then compare the empirical performance of reinforcement learning with verifiable rewards in post-training using both GRPO variants. Our results show that off-policy GRPO either significantly outperforms or performs on par with its on-policy counterpart.

强化学习策略优化离线训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。