arXiv:2603.14041cs.AI2026-03

用反思奖励提升大模型数学推理能力,效果优于传统方法。

GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models

  • 引入分阶段框架,结合群体相对策略优化与反思奖励。
  • 在数学推理任务中达到当前最佳表现,反思奖励至关重要。
  • 适合研究大模型自我反思机制或智能体训练的学者参考。

大型语言模型(LLMs)推理能力的提升受到广泛关注,监督微调(SFT)和强化学习成为主流范式。尽管已有研究认识到反思在推理过程中的重要性,但现有方法极少主动鼓励训练中的反思行为。本文聚焦数学推理,提出一种四阶段框架,融合群体相对策略优化(GRPO)与反思奖励机制,以增强模型的自省能力。该方法还整合了既有的准确率与格式奖励。实验表明,通过鼓励反思的训练,GRPO实现领先性能;消融实验证实反思奖励的关键作用。对比评估显示,全参数SFT优于低秩适配(LoRA),尽管计算成本更高。基于这些发现,本研究证实GRPO在后训练优化中的方法论价值,并展望其作为未来基于大模型智能体的核心驱动力,通过认知奖励与动态环境交互的协同,实现突破。

原文摘要 · Abstract (English)

The enhancement of reasoning capabilities in large language models (LLMs) has garnered significant attention, with supervised fine-tuning (SFT) and reinforcement learning emerging as dominant paradigms. While recent studies recognize the importance of reflection in reasoning processes, existing methodologies seldom address proactive reflection encouragement during training. This study focuses on mathematical reasoning by proposing a four-stage framework integrating Group Relative Policy Optimization (GRPO) with reflection reward mechanisms to strengthen LLMs' self-reflective capabilities. Besides, this approach incorporates established accuracy and format reward. Experimental results demonstrate GRPO's state-of-the-art performance through reflection-encouraged training, with ablation studies confirming the reflection reward's pivotal role. Comparative evaluations demonstrate full-parameter SFT's superiority over low-rank adaptation (LoRA) despite heightened computational demands. Building on these cumulative findings, this research substantiates GRPO's methodological significance in post-training optimization and envisions its potential to serve as a pivotal enabler for future LLM-based intelligent agents through the synergistic integration of cognitive rewards with dynamic environmental interactions.

大模型数学推理强化学习反思机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。