arXiv:2501.07301cs.CLcs.AI2025-01ACL被引 468

改进数学推理中的过程奖励模型,提升中间步骤纠错能力。

The Lessons of Developing Process Reward Models in Mathematical Reasoning

  • 用大模型自评与人工标注结合,替代传统蒙特卡洛估算方法。
  • 发现现有评估策略导致模型关注答案而非推理过程,影响准确性。
  • 提出共识过滤机制,显著提升模型在步骤级错误识别上的表现。

过程奖励模型(PRM)为大型语言模型在数学推理中的过程监督提供了新路径,旨在识别并纠正推理过程中的中间错误。然而,其开发面临数据标注与评估方法的重大挑战。本文通过大量实验表明,常见的基于蒙特卡洛(MC)估计的数据合成方法,相比大模型自评和人工标注,在性能与泛化能力上均表现更差。这是因为MC依赖生成模型判断当前步骤正确性,易造成验证不准。此外,我们发现传统最佳选一(BoN)评估策略存在三大问题:(1)不可靠的策略模型虽得正确答案但过程有误,导致评估标准与PRM目标不一致;(2)PRM对这类响应容忍度高,导致BoN得分虚高;(3)现有PRM中大量最低分集中在最终答案步骤,表明其已从过程导向转向结果导向。为此,我们设计了一种共识过滤机制,有效融合MC估计与大模型自评,并倡导结合响应级与步骤级指标的综合评估框架。基于此,我们在BoN评估与步骤级错误识别任务中显著提升了模型表现与数据效率。最后,我们发布了一个新状态最优的PRM,超越现有开源模型,并提供未来构建过程监督模型的实用指南。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) emerge as a promising approach for process supervision in mathematical reasoning of Large Language Models (LLMs), which aim to identify and mitigate intermediate errors in the reasoning processes. However, the development of effective PRMs faces significant challenges, particularly in data annotation and evaluation methodologies. In this paper, through extensive experiments, we demonstrate that commonly used Monte Carlo (MC) estimation-based data synthesis for PRMs typically yields inferior performance and generalization compared to LLM-as-a-judge and human annotation methods. MC estimation relies on completion models to evaluate current-step correctness, leading to inaccurate step verification. Furthermore, we identify potential biases in conventional Best-of-N (BoN) evaluation strategies for PRMs: (1) The unreliable policy models generate responses with correct answers but flawed processes, leading to a misalignment between the evaluation criteria of BoN and the PRM objectives of process verification. (2) The tolerance of PRMs of such responses leads to inflated BoN scores. (3) Existing PRMs have a significant proportion of minimum scores concentrated on the final answer steps, revealing the shift from process to outcome-based assessment in BoN Optimized PRMs. To address these challenges, we develop a consensus filtering mechanism that effectively integrates MC estimation with LLM-as-a-judge and advocates a more comprehensive evaluation framework that combines response-level and step-level metrics. Based on the mechanisms, we significantly improve both model performance and data efficiency in the BoN evaluation and the step-wise error identification task. Finally, we release a new state-of-the-art PRM that outperforms existing open-source alternatives and provides practical guidelines for future research in building process supervision models.

数学推理过程奖励大模型评估奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。