arXiv:2511.12344cs.AI2025-11被引 20

用评分标准提升大模型多领域推理能力,探索更广解空间。

Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning

  • 引入评分标准提供细粒度奖励与离线指导,扩展探索范围。
  • 在14个跨领域任务中平均提升7%-8.4%,优于纯在线强化学习。
  • 适合需要持续探索突破性能瓶颈的复杂推理场景。

近期强化学习进展显著提升了大语言模型(LLMs)的复杂推理能力。然而,现有方法主要聚焦于单一领域(如数学)的可验证奖励强化学习(RLVR),且依赖纯在线框架,限制了探索空间,影响推理表现。本文提出一种基于评分标准的强化学习框架RGR-GRPO,通过评分标准提供细粒度奖励信号与离线引导,使模型在GRPO训练中获得更密集、更丰富的奖励,并拓展解空间。在14个跨领域基准测试中,RGR-GRPO持续优于仅依赖其他奖励机制或离线引导的强化学习方法。相比可验证在线强化学习基线,在数学、物理、化学和通用推理任务上分别实现+7.0%、+5.4%、+8.4%和+6.6%的平均提升。值得注意的是,RGR-GRPO在离策略训练中保持稳定的熵波动,展现出持续探索能力,有效突破现有性能瓶颈。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning (RL) have significantly improved the complex reasoning capabilities of large language models (LLMs). Despite these successes, existing methods mainly focus on single-domain RL (e.g., mathematics) with verifiable rewards (RLVR), and their reliance on purely online RL frameworks restricts the exploration space, thereby limiting reasoning performance. In this paper, we address these limitations by leveraging rubrics to provide both fine-grained reward signals and offline guidance. We propose $\textbf{RGR-GRPO}$ (Reward and Guidance through Rubrics), a rubric-driven RL framework for multi-domain reasoning. RGR-GRPO enables LLMs to receive dense and informative rewards while exploring a larger solution space during GRPO training. Extensive experiments across 14 benchmarks spanning multiple domains demonstrate that RGR-GRPO consistently outperforms RL methods that rely solely on alternative reward schemes or offline guidance. Compared with verifiable online RL baseline, RGR-GRPO achieves average improvements of +7.0%, +5.4%, +8.4%, and +6.6% on mathematics, physics, chemistry, and general reasoning tasks, respectively. Notably, RGR-GRPO maintains stable entropy fluctuations during off-policy training and achieves superior pass@k performance, reflecting sustained exploration and effective breakthrough beyond existing performance bottlenecks.

强化学习多领域推理评分标准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。