arXiv:2506.17533cs.CL2025-06被引 2

用双重奖励提升大模型数学推理能力

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning

  • 同时学习步骤正确性与最终答案潜力
  • 在MATH500和ProcessBench上超越单一奖励模型
  • 适合需要高精度数学推理的场景

本文提出DuaShepherd,一种融合步骤正确性与最终答案潜力双重奖励信号的新型奖励建模框架,以增强大语言模型(LLMs)的数学推理能力。正确性信号关注步骤中的错误识别,潜力信号则聚焦于达成正确答案的可能性。我们构建了包含双信号的大规模自动化奖励数据集,并采用统一多头架构,在多任务设置下联合训练两个奖励模型,实现并行学习的优势。通过将两种信号融合为复合概率,模型在多个基准测试中均取得稳定提升。在MATH500和ProcessBench上的实证评估表明,该联合奖励显著优于仅使用任一奖励类型训练的模型,在相近资源约束下达到当前最优性能。

原文摘要 · Abstract (English)

In this paper, we propose DuaShepherd, a novel reward modeling framework that integrates two complementary reward signals, correctness and potential, to enhance the mathematical reasoning capabilities of Large Language Models (LLMs). While correctness-based signals emphasize identification of stepwise errors, potential-based signals focus on the likelihood of reaching the correct final answer. We developed an automated pipeline for constructing large-scale reward modeling dataset with both signals. A unified, multi-head architecture was explored to train the two reward models in a multi-task setup, demonstrating benefits from learning both correctness and potential in parallel. By combining these two signals into a compound probability, our model achieves consistent performance improvements across multiple benchmarks. Empirical evaluations on MATH500 and ProcessBench confirm that this combined reward significantly outperforms models trained on either reward type alone, achieving state-of-the-art performance under comparable resource constraints.

数学推理奖励建模大模型多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。