arXiv:2503.21295cs.CL2025-03EMNLP被引 48

用推理驱动方法提升数学推理评分模型的准确性和泛化能力。

R-PRM: Reasoning-Driven Process Reward Modeling

  • 用强模型生成标注数据,弥补标注稀缺问题
  • 通过偏好优化提升性能,无需额外标注
  • 推理时扩展能力,显著提升六大数据集准确率

大型语言模型在进行逐步数学推理时不可避免会出错。过程奖励模型(PRMs)通过评估每一步推理成为有前景的解决方案。然而,现有PRMs通常直接输出评分,限制了学习效率和评估准确性,且受限于标注数据稀缺。为此,我们提出推理驱动的过程奖励建模(R-PRM)。首先,利用更强的LLM从有限标注中生成种子数据,有效提升模型推理能力并实现全面的步骤评估。其次,通过偏好优化进一步提升性能,无需额外标注。第三,引入推理时缩放以充分释放模型推理潜力。大量实验表明,R-PRM在ProcessBench和PRMBench上分别超越强基线11.9和8.5个F1分数;应用于引导数学推理时,在六个挑战性数据集上一致提升超过8.5个百分点。进一步分析显示,R-PRM具有更全面的评估能力和更强的泛化能力,凸显其巨大潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by evaluating each reasoning step. However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy, which is further exacerbated by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM). First, we leverage stronger LLMs to generate seed data from limited annotations, effectively bootstrapping our model's reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we further enhance performance through preference optimization, without requiring additional annotated data. Third, we introduce inference-time scaling to fully harness the model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 11.9 and 8.5 points in F1 scores, respectively. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.5 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and stronger generalization capabilities, thereby highlighting its significant potential.

推理建模奖励模型数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。