arXiv:2503.02382cs.CLcs.AI2025-03ACL被引 18

用自适应搜索高效构建高质量数学推理步骤标注数据

An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning

  • 基于步骤贡献度量化自动标注中间推理过程
  • 构建5万条高质量标注数据集Epic50k,提升奖励模型性能
  • 适合需要高精度数学推理训练数据的研究者使用

提升大语言模型的数学推理能力具有重要的科学与实际意义。研究者通常采用过程监督奖励模型(PRM)来引导推理过程,有效增强模型的推理能力。然而,现有构建过程监督训练数据的方法,如人工标注和逐步蒙特卡洛估计,往往成本高昂或质量不佳。为此,本文提出名为EpicPRM的框架,基于每一步推理的量化贡献进行标注,并采用自适应二分搜索算法,显著提升标注的精度与效率。利用该方法,我们高效构建了一个名为Epic50k的高质量过程监督训练数据集,包含5万条标注的中间推理步骤。相比其他公开数据集,基于Epic50k训练的PRM表现出显著更优的性能。Epic50k数据集获取地址:https://github.com/xiaolizh1/EpicPRM。

原文摘要 · Abstract (English)

Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised reward models (PRMs) to guide the reasoning process, effectively improving the models' reasoning abilities. However, existing methods for constructing process supervision training data, such as manual annotation and per-step Monte Carlo estimation, are often costly or suffer from poor quality. To address these challenges, this paper introduces a framework called EpicPRM, which annotates each intermediate reasoning step based on its quantified contribution and uses an adaptive binary search algorithm to enhance both annotation precision and efficiency. Using this approach, we efficiently construct a high-quality process supervision training dataset named Epic50k, consisting of 50k annotated intermediate steps. Compared to other publicly available datasets, the PRM trained on Epic50k demonstrates significantly superior performance. Getting Epic50k at https://github.com/xiaolizh1/EpicPRM.

数学推理奖励模型数据构建LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。