新标注方法让模型更懂数学推理中的自我修正。
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
- 基于错误传播与终止概念,重新标注长链思维过程。
- 在170万样本上训练的7B模型,多项指标超越现有方法。
- 适合需要精准评估推理过程的研究者与开发者。
现有方法在处理长链思维(CoT)推理时,仅关注首个错误步骤及之前的内容,忽略后续可能存在的正确推理。为解决此问题,本文提出一种针对长CoT推理过程的新数据标注方法,引入‘错误传播’与‘错误终止’概念,以捕捉反思模式中正确与错误步骤交替的现象。利用大模型进行标注,构建了包含170万样本的数据集,用于训练一个70亿参数的PRM,并在解题与步骤层面进行评估。实验表明,该模型在搜索引导、BoN和F1等指标上均优于现有开源PRM及基于MC的标注方法。详细分析验证了方法的稳定性与泛化能力。
原文摘要 · Abstract (English)
Many studies focus on data annotation techniques for training effective PRMs. However, current methods encounter a significant issue when applied to long CoT reasoning processes: they tend to focus solely on the first incorrect step and all preceding steps, assuming that all subsequent steps are incorrect. These methods overlook the unique self-correction and reflection mechanisms inherent in long CoT, where correct reasoning steps may still occur after initial reasoning mistakes. To address this issue, we propose a novel data annotation method for PRMs specifically designed to score the long CoT reasoning process. Given that under the reflection pattern, correct and incorrect steps often alternate, we introduce the concepts of Error Propagation and Error Cessation, enhancing PRMs' ability to identify both effective self-correction behaviors and reasoning based on erroneous steps. Leveraging an LLM-based judger for annotation, we collect 1.7 million data samples to train a 7B PRM and evaluate it at both solution and step levels. Experimental results demonstrate that compared to existing open-source PRMs and PRMs trained on open-source datasets, our PRM achieves superior performance across various metrics, including search guidance, BoN, and F1 scores. Compared to widely used MC-based annotation methods, our annotation approach not only achieves higher data efficiency but also delivers superior performance. Detailed analysis is also conducted to demonstrate the stability and generalizability of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。