系统梳理奖励模型与学习策略,揭示大模型如何通过反馈实现智能进化
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- 从训练到推理全流程解析奖励驱动学习机制
- 涵盖RLHF、DPO等主流方法,支持动态反馈与偏好对齐
- 适合研究者快速掌握奖励学习全景,含开源论文合集
大型语言模型(LLMs)的发展已从预训练扩展转向后训练与测试时扩展。在此背景下,一种核心统一范式——基于奖励的学习(Learning from Rewards)逐渐成型,其中奖励信号作为引导星,指导大模型行为。该范式支撑了包括强化学习(RLHF、RLAIF、DPO、GRPO)、奖励引导解码及事后修正在内的多种主流技术。关键在于,它实现了从被动学习静态数据到主动学习动态反馈的转变,使大模型具备任务适配的偏好对齐与深度推理能力。本文从奖励模型与学习策略角度,系统综述了训练、推理及推理后阶段的奖励学习方法,讨论了奖励模型评估基准与主要应用场景,并指出当前挑战与未来方向。相关论文集合见:https://github.com/bobxwu/learning-from-rewards-llm-papers。
原文摘要 · Abstract (English)
Recent developments in Large Language Models (LLMs) have shifted from pre-training scaling to post-training and test-time scaling. Across these developments, a key unified paradigm has arisen: Learning from Rewards, where reward signals act as the guiding stars to steer LLM behavior. It has underpinned a wide range of prevalent techniques, such as reinforcement learning (RLHF, RLAIF, DPO, and GRPO), reward-guided decoding, and post-hoc correction. Crucially, this paradigm enables the transition from passive learning from static data to active learning from dynamic feedback. This endows LLMs with aligned preferences and deep reasoning capabilities for diverse tasks. In this survey, we present a comprehensive overview of learning from rewards, from the perspective of reward models and learning strategies across training, inference, and post-inference stages. We further discuss the benchmarks for reward models and the primary applications. Finally we highlight the challenges and future directions. We maintain a paper collection at https://github.com/bobxwu/learning-from-rewards-llm-papers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。