arXiv:2508.00222cs.AIcs.CL2025-08ACL被引 38

解决大模型强化学习中能力边界坍缩问题,提升推理能力

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

  • 融合内部探索与外部数据的混合策略优化方法
  • 在6个数学推理基准上达顶尖性能,平均提升69.2%
  • 有效突破基座模型能力限制,适合需要强推理的场景

基于可验证奖励的强化学习(RLVR)显著提升了大语言模型(LLM)的复杂推理能力。然而,由于其本质上是单策略方法,加之LLM动作空间巨大且奖励稀疏,RLVR难以突破基座LLM的固有能力边界,甚至导致能力边界坍缩,缩小模型解题范围。为此,我们提出RL-PLUS,一种针对LLM的新型混合策略优化方法,通过内生利用与外部数据协同,实现更强的推理能力并超越基座模型边界。RL-PLUS包含两大核心组件:多重重要性采样以缓解外部数据分布偏差,以及基于探索的优势函数,引导模型走向高价值、未探索的推理路径。我们提供了理论分析与大量实验,验证该方法的优越性与泛化性。相比现有RLVR方法,RL-PLUS在六个数学推理基准上达到领先水平;在六个分布外推理任务中表现更优;在多种模型家族中均实现持续显著提升,平均相对改进高达69.2%。此外,Pass@k曲线分析表明,RL-PLUS有效解决了能力边界坍缩问题。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs). However, it struggles to break through the inherent capability boundaries of the base LLM, due to its essentially on-policy strategy coupled with LLM's immense action space and sparse reward. Critically, RLVR can lead to the capability boundary collapse, narrowing the LLM's problem-solving scope. To address this problem, we propose RL-PLUS, a novel hybrid-policy optimization approach for LLMs that synergizes internal exploitation with external data to achieve stronger reasoning capabilities and surpass the boundaries of base models. RL-PLUS integrates two core components, i.e., Multiple Importance Sampling to address distributional mismatch from external data, and Exploration-Based Advantage Function to guide the model towards high-value, unexplored reasoning paths. We provide both theoretical analysis and extensive experiments to demonstrate the superiority and generalizability of our approach. Compared with existing RLVR methods, RL-PLUS achieves 1) state-of-the-art performance on six math reasoning benchmarks; 2) superior performance on six out-of-distribution reasoning tasks; 3) consistent and significant gains across diverse model families, with average relative improvements up to 69.2\%. Moreover, the analysis of Pass@k curves indicates that RL-PLUS effectively resolves the capability boundary collapse problem.

强化学习大模型推理混合策略数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。