arXiv:2510.02245cs.LGcs.AI2025-10被引 46

通过分析推理经验价值,提升大模型强化学习效率与稳定性。

ExGRPO: Learning to Reason from Experience

  • 基于推理正确性和熵值筛选高价值经验,动态管理训练数据。
  • 在5个模型上平均提升数学/通用推理成绩3.5/7.6分,超越传统方法。
  • 适合研究高效强化学习或大模型推理优化的学者与工程师。

从可验证奖励中进行强化学习(RLVR)是提升大语言模型推理能力的新范式。然而,标准的在线策略训练在单次更新后丢弃回溯经验,导致计算效率低且训练不稳定。尽管过往强化学习研究已表明复用历史经验的优势,但经验特性如何影响大推理模型的学习动态仍缺乏探索。本文首次揭示了回溯正确性与熵值是衡量经验价值的有效指标。据此提出ExGRPO(经验组相对策略优化)框架,通过组织与优先排序高价值经验,并采用混合策略目标平衡探索与经验利用。在五个主干模型(1.5B–8B参数)上的实验表明,ExGRPO在数学与通用基准上均持续提升推理性能,平均较在线策略RLVR提升+3.5/+7.6分;同时在强弱模型上均稳定训练,而在线方法在此类场景下失效。结果凸显了有原则的经验管理是实现高效、可扩展RLVR的关键。

原文摘要 · Abstract (English)

Reinforcement learning from verifiable rewards (RLVR) is an emerging paradigm for improving the reasoning ability of large language models. However, standard on-policy training discards rollout experiences after a single update, leading to computational inefficiency and instability. While prior work on RL has highlighted the benefits of reusing past experience, the role of experience characteristics in shaping learning dynamics of large reasoning models remains underexplored. In this paper, we are the first to investigate what makes a reasoning experience valuable and identify rollout correctness and entropy as effective indicators of experience value. Based on these insights, we propose ExGRPO (Experiential Group Relative Policy Optimization), a framework that organizes and prioritizes valuable experiences, and employs a mixed-policy objective to balance exploration with experience exploitation. Experiments on five backbone models (1.5B-8B parameters) show that ExGRPO consistently improves reasoning performance on mathematical/general benchmarks, with an average gain of +3.5/7.6 points over on-policy RLVR. Moreover, ExGRPO stabilizes training on both stronger and weaker models where on-policy methods fail. These results highlight principled experience management as a key ingredient for efficient and scalable RLVR.

强化学习推理优化经验管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。