arXiv:2412.11006cs.LGcs.CL2024-12被引 28

用熵正则化提升大模型数学推理的步骤奖励效果

Entropy-Regularized Process Reward Model

  • 引入熵正则项约束策略分布,防止偏离初始行为
  • 在GSM8K和MATH上比现有方法提升1%-3%
  • 适合需要精准推理轨迹的复杂任务研究者

大语言模型在多步推理中表现优异,但在数学推理方面仍存在系统性错误。一种有前景的解决方案是通过奖励模型引导强化学习,尤其是关注过程奖励的模型,其对每个中间步骤进行评分,而非仅评估最终结果。本文提出熵正则化过程奖励模型(ER-PRM),结合KL正则化的马尔可夫决策过程,平衡策略优化与避免策略过度偏离初始分布的需求。我们基于理论推导出一种新的奖励构建方法,证明可从初始策略采样中获得最优奖励模型。在MATH和GSM8K基准上的实验证明,ER-PRM持续优于现有过程奖励模型,在最佳N种方案评估下,于GSM8K提升1%,于MATH提升2-3%;在RLHF下提升超过1%。这些结果凸显了熵正则化在增强大模型推理能力方面的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic errors. A promising solution is reinforcement learning (RL) guided by reward models, particularly those focusing on process rewards, which score each intermediate step rather than solely evaluating the final outcome. This approach is more effective at guiding policy models towards correct reasoning trajectories. In this work, we propose an entropy-regularized process reward model (ER-PRM) that integrates KL-regularized Markov Decision Processes (MDP) to balance policy optimization with the need to prevent the policy from shifting too far from its initial distribution. We derive a novel reward construction method based on the theoretical results. Our theoretical analysis shows that we could derive the optimal reward model from the initial policy sampling. Our empirical experiments on the MATH and GSM8K benchmarks demonstrate that ER-PRM consistently outperforms existing process reward models, achieving 1% improvement on GSM8K and 2-3% improvement on MATH under best-of-N evaluation, and more than 1% improvement under RLHF. These results highlight the efficacy of entropy-regularization in enhancing LLMs' reasoning capabilities.

推理增强强化学习奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。