arXiv:2510.08049cs.CLcs.AI2025-10综述被引 31

让大模型学会一步步思考,比只看答案更准

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

  • 用中间推理步骤替代最终答案来训练奖励模型
  • 覆盖数学、代码、多模态等多领域应用
  • 适合研究大模型推理对齐与强化学习的人

尽管大语言模型具备强大推理能力,传统对齐方法仍主要依赖仅评价最终答案的结果奖励模型(ORMs)。过程奖励模型(PRMs)通过在步骤或轨迹层面评估和引导推理,弥补了这一不足。本综述系统梳理了PRMs的完整流程:如何生成过程数据、构建PRMs,以及如何用于测试时扩展和强化学习。我们总结了其在数学、编程、文本、多模态推理、机器人和智能体等领域的应用,并回顾了新兴基准。目标是厘清设计空间,揭示开放挑战,引导未来研究向细粒度、鲁棒的推理对齐发展。

原文摘要 · Abstract (English)

Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) address this gap by evaluating and guiding reasoning at the step or trajectory level. This survey provides a systematic overview of PRMs through the full loop: how to generate process data, build PRMs, and use PRMs for test-time scaling and reinforcement learning. We summarize applications across math, code, text, multimodal reasoning, robotics, and agents, and review emerging benchmarks. Our goal is to clarify design spaces, reveal open challenges, and guide future research toward fine-grained, robust reasoning alignment.

大模型推理对齐奖励模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。