arXiv:2507.18071cs.LGcs.AI2025-07被引 657

GSPO通过序列级优化提升大模型强化学习训练效率与稳定性。

Group Sequence Policy Optimization

  • 基于序列似然定义重要性比率,进行序列级裁剪与优化
  • 相比GRPO显著提升训练效率,稳定MoE强化学习训练
  • 适合需要高效稳定强化学习的大模型研发团队

本文提出群组序列策略优化(GSPO),一种稳定、高效且高性能的强化学习算法,用于训练大语言模型。与以往基于词元级重要性比率的算法不同,GSPO基于序列似然定义重要性比率,并执行序列级裁剪、奖励与优化。实验表明,相较于GRPO算法,GSPO在训练效率和性能上均表现更优,尤其能有效稳定混合专家(MoE)强化学习训练过程,并有潜力简化强化学习基础设施设计。这些优势已助力最新版Qwen3模型实现显著性能提升。

原文摘要 · Abstract (English)

This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models. Unlike previous algorithms that adopt token-level importance ratios, GSPO defines the importance ratio based on sequence likelihood and performs sequence-level clipping, rewarding, and optimization. We demonstrate that GSPO achieves superior training efficiency and performance compared to the GRPO algorithm, notably stabilizes Mixture-of-Experts (MoE) RL training, and has the potential for simplifying the design of RL infrastructure. These merits of GSPO have contributed to the remarkable improvements in the latest Qwen3 models.

强化学习大模型训练序列优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。