arXiv:2508.03556cs.LG2025-08被引 9

用视觉推理提升语言模型的思维评估能力,大幅降低标注成本

VRPRM: Process Reward Modeling via Visual Reasoning

  • 通过两阶段训练融合视觉推理,实现高效思维评估
  • 仅需3.6K标注数据即超越400K数据训练的基线模型
  • 适合需要低成本高质量思维评估的AI系统研发

过程奖励模型(PRM)广泛用于大语言模型的后训练阶段,因其能对生成内容的推理步骤进行细粒度评估。然而,现有多数PRM缺乏长期推理与深度思考能力。尽管已有研究尝试将思维链(CoT)引入PRM,但其标注成本过高,难以在各类任务中稳定应用。为此,本文提出基于视觉推理的过程奖励模型VRPRM,并设计高效的两阶段训练策略。实验表明,仅使用3.6K CoT-PRM监督微调数据和50K非CoT PRM强化学习数据,VRPRM即可超越总数据量达400K的非思考型PRM,在BoN测试中相对于基线模型性能提升最高达118%。结果证实,该联合训练策略能在更低标注成本下实现更优的推理能力,为PRM训练提供了更高效的数据利用新范式。

原文摘要 · Abstract (English)

Process Reward Model (PRM) is widely used in the post-training of Large Language Model (LLM) because it can perform fine-grained evaluation of the reasoning steps of generated content. However, most PRMs lack long-term reasoning and deep thinking capabilities. On the other hand, although a few works have tried to introduce Chain-of-Thought (CoT) capability into PRMs, the annotation cost of CoT-PRM data is too expensive to play a stable role in various tasks. To address the above challenges, we propose VRPRM, a process reward model via visual reasoning, and design an efficient two-stage training strategy. Experimental results show that using only 3.6K CoT-PRM Supervised Fine-Tuning(SFT) data and 50K non-CoT PRM Reinforcement Learning (RL) training data, VRPRM can surpass the non-thinking PRM with a total data volume of 400K and achieved a relative performance improvement of up to 118\% over the base model in the BoN experiment. This result confirms that the proposed combined training strategy can achieve higher quality reasoning capabilities at a lower data annotation cost, thus providing a new paradigm for PRM training with more efficient data utilization.

奖励建模思维链低资源训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。