arXiv:2510.01925cs.CL2025-10中稿 · publication in Art…综述被引 6

用奖励模型提升大模型推理能力,系统梳理方法与应用。

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

  • 构建奖励模型捕捉人类偏好,用于优化生成与选择最佳答案。
  • 支持推理过程中的自动纠错与自我改进,显著提升回答质量。
  • 适合研究强化学习增强大模型、构建可信生成系统的开发者。

奖励模型(RMs)在提升大语言模型(LLM)推理能力中起关键作用。它们可为强化学习微调提供训练信号,并在推理阶段从多个候选答案中选出最优解。本文系统介绍奖励模型的基本概念,涵盖其架构设计、训练方法与评估技术。进一步探讨其三大核心应用:(1)指导生成与推理阶段的最优输出选择;(2)促进数据合成与模型迭代自提升;(3)在基于强化学习的微调中提供训练信号。最后,结合现有研究与实证发现,讨论奖励模型在选择、泛化、评估与增强方面的关键开放问题。本分析旨在为奖励模型在大模型推理中的有效部署与持续发展提供可操作的洞见。

原文摘要 · Abstract (English)

Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learning (RL) and help select the best answer from multiple candidates during inference. In this paper, we provide a systematic introduction to RMs, along with a comprehensive survey of their applications in LLM reasoning. We first review fundamental concepts of RMs, including their architectures, training methodologies, and evaluation techniques. Then, we explore their key applications: (1) guiding generation and selecting optimal outputs during LLM inference, (2) facilitating data synthesis and iterative self-improvement for LLMs, and (3) providing training signals in RL-based finetuning. Finally, we discuss critical open questions regarding the selection, generalization, evaluation, and enhancement of RMs, based on existing research and our own empirical findings. Our analysis aims to provide actionable insights for the effective deployment and advancement of RMs for LLM reasoning.

奖励模型大模型推理强化学习自提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。