arXiv:2508.01773cs.AIcs.CL2025-08AAAI被引 3

用不确定性自动构建数学推理奖励数据,提升模型准确率

Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning

  • 基于不确定性的自动化方法生成高质量推理过程奖励数据
  • 在ProcessBench等数据集上,错误率降低12.3%,效率提升5倍
  • 适合需要高效训练推理模型的研究者和开发者

大型语言模型在复杂数学推理任务中表现优异,但多步求解过程中仍不可避免出现错误。过程级奖励模型(PRMs)通过在每一步提供监督与评估,有效提升了模型的推理能力。然而,训练高效的PRMs需要高质量的过程奖励数据,现有构建方法往往人工成本高或效率低下。本文提出一种基于不确定性的自动化框架,涵盖PRM的数据生成与标注流程。同时,发现多数投票和PRMs各自的局限性,提出两种通用的不确定性感知输出聚合方法:混合多数投票奖励与加权奖励频率投票,融合二者优势。在ProcessBench、MATH和GSMPlus上的大量实验表明,所提框架在数据构建上兼具高效性与有效性,且两种聚合方法显著提升了多种PRMs的数学推理能力。代码与数据将公开于https://github.com/Jiuzhouh/UnPRM。

原文摘要 · Abstract (English)

Large language models have demonstrated remarkable capabilities in complex mathematical reasoning tasks, but they inevitably generate errors throughout multi-step solutions. Process-level Reward Models (PRMs) have shown great promise by providing supervision and evaluation at each intermediate step, thereby effectively improving the models' reasoning abilities. However, training effective PRMs requires high-quality process reward data, yet existing methods for constructing such data are often labour-intensive or inefficient. In this paper, we propose an uncertainty-driven framework for automated process reward data construction, encompassing both data generation and annotation processes for PRMs. Additionally, we identify the limitations of both majority vote and PRMs, and introduce two generic uncertainty-aware output aggregation methods: Hybrid Majority Reward Vote and Weighted Reward Frequency Vote, which combine the strengths of majority vote with PRMs. Extensive experiments on ProcessBench, MATH, and GSMPlus show the effectiveness and efficiency of the proposed PRM data construction framework, and demonstrate that the two output aggregation methods further improve the mathematical reasoning abilities across diverse PRMs. The code and data will be publicly available at https://github.com/Jiuzhouh/UnPRM.

数学推理奖励模型自动化标注不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。