提出LUSPO算法,解决强化学习中响应长度偏差问题
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
- 修正GSPO的长度偏差,使损失函数与响应长度无关
- 在数学和多模态推理任务中性能超越GRPO、GSPO
- 适合需要稳定生成长度的复杂推理场景
最近基于可验证奖励的强化学习(RLVR)在大语言模型(LLMs)和视觉-语言模型(VLMs)中的应用,显著提升了复杂任务的推理能力。训练过程中,响应长度增长常被视为推理能力提升的关键因素。然而,不同RLVR算法在训练中响应长度的变化模式差异显著。本文深入分析主流RLVR算法的组成,提出理论解释响应长度变化的影响因素,并通过大量实验验证。基于此,提出长度无偏序列策略优化(LUSPO)算法,修正了组序列策略优化(GSPO)固有的长度偏差,使其损失函数对响应长度无偏,从而解决响应长度坍缩问题。在数学推理基准和多模态推理场景中,LUSPO均实现更优性能,实证表明其是当前最先进的优化策略,优于GRPO和GSPO。
原文摘要 · Abstract (English)
Recent applications of Reinforcement Learning with Verifiable Rewards (RLVR) to Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated significant success in enhancing reasoning capabilities for complex tasks. During RLVR training, an increase in response length is often regarded as a key factor contributing to the growth of reasoning ability. However, the patterns of change in response length vary significantly across different RLVR algorithms during the training process. To provide a fundamental explanation for these variations, this paper conducts an in-depth analysis of the components of mainstream RLVR algorithms. We present a theoretical analysis of the factors influencing response length and validate our theory through extensive experimentation. Building upon these theoretical findings, we propose the Length-Unbiased Sequence Policy Optimization (LUSPO) algorithm. Specifically, we rectify the length bias inherent in Group Sequence Policy Optimization (GSPO), rendering its loss function unbiased with respect to response length and thereby resolving the issue of response length collapse. We conduct extensive experiments across mathematical reasoning benchmarks and multimodal reasoning scenarios, where LUSPO consistently achieves superior performance. Empirical results demonstrate that LUSPO represents a novel, state-of-the-art optimization strategy compared to existing methods such as GRPO and GSPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。