arXiv:2506.03570cs.CL2025-06被引 11

无需真实步骤标签,用结果正确性生成伪标签训练过程奖励模型。

FreePRM: Training Process Reward Models Without Ground Truth Process Labels

  • 根据最终结果正确性生成伪步骤标签,降低标注成本。
  • 在ProcessBench上达53.0% F1,超越监督模型24.1%。
  • 适合无大量人工标注数据的研究者使用。

大型语言模型的进展表明,过程奖励模型(PRMs)对提升模型性能至关重要。然而,训练PRMs通常需要逐步标签,无论是人工标注还是自动生成,都存在成本高、难以规模化的问题。为此,本文提出FreePRM,一种无需真实步骤标签的弱监督训练框架。该方法首先基于最终结果的正确性生成伪步骤标签,再利用缓冲概率机制消除伪标签中的噪声影响。实验表明,FreePRM在ProcessBench上的平均F1得分为53.0%,比在Math-Shepherd上训练的全监督PRM高出24.1%。相较于其他开源PRM,其表现优于RLHFlow-PRM-Mistral-8B(28.4%)24.6%、EurusPRM(31.3%)21.7%、Skywork-PRM-7B(42.1%)10.9%。本工作提出了一种新的PRM训练范式,显著降低对昂贵步骤标注的依赖,同时保持优异性能。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have demonstrated that Process Reward Models (PRMs) play a crucial role in enhancing model performance. However, training PRMs typically requires step-level labels, either manually annotated or automatically generated, which can be costly and difficult to obtain at scale. To address this challenge, we introduce FreePRM, a weakly supervised framework for training PRMs without access to ground-truth step-level labels. FreePRM first generates pseudo step-level labels based on the correctness of final outcome, and then employs Buffer Probability to eliminate impact of noise inherent in pseudo labeling. Experimental results show that FreePRM achieves an average F1 score of 53.0% on ProcessBench, outperforming fully supervised PRM trained on Math-Shepherd by +24.1%. Compared to other open-source PRMs, FreePRM outperforms upon RLHFlow-PRM-Mistral-8B (28.4%) by +24.6%, EurusPRM (31.3%) by +21.7%, and Skywork-PRM-7B (42.1%) by +10.9%. This work introduces a new paradigm in PRM training, significantly reducing reliance on costly step-level annotations while maintaining strong performance.

奖励模型弱监督过程评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。