强化学习让大模型更懂人意、推理更强,这篇综述系统梳理了全过程应用。
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- 用强化学习提升大模型从预训练到推理的全周期能力
- 验证性奖励机制显著增强模型推理表现
- 适合关注大模型智能升级的研究者与开发者
近年来,以强化学习(RL)为核心的训练方法显著提升了大语言模型(LLMs)在理解人类意图、遵循指令和增强推理能力方面的表现。尽管已有综述对RL增强的LLM进行了概述,但多数范围有限,未能涵盖RL在LLM全生命周期中的作用。本文系统回顾了RL赋能LLM的理论与实践进展,尤其聚焦于可验证奖励的强化学习(RLVR)。首先简要介绍RL基础理论;其次详细阐述RL在预训练、对齐微调和强化推理等各阶段的应用策略,特别指出强化推理阶段是推动模型推理能力突破的关键动力。接着,整理了当前用于RL微调的主流数据集与评估基准,涵盖人工标注数据、AI辅助偏好数据及程序验证类语料。随后,综述主流开源工具与训练框架,为后续研究提供实用参考。最后分析该领域未来挑战与趋势。本综述旨在为研究者与从业者呈现RL与LLM交叉领域的最新进展与前沿方向,助力打造更智能、通用且安全的大模型。
原文摘要 · Abstract (English)
In recent years, training methods centered on Reinforcement Learning (RL) have markedly enhanced the reasoning and alignment performance of Large Language Models (LLMs), particularly in understanding human intents, following user instructions, and bolstering inferential strength. Although existing surveys offer overviews of RL augmented LLMs, their scope is often limited, failing to provide a comprehensive summary of how RL operates across the full lifecycle of LLMs. We systematically review the theoretical and practical advancements whereby RL empowers LLMs, especially Reinforcement Learning with Verifiable Rewards (RLVR). First, we briefly introduce the basic theory of RL. Second, we thoroughly detail application strategies for RL across various phases of the LLM lifecycle, including pre-training, alignment fine-tuning, and reinforced reasoning. In particular, we emphasize that RL methods in the reinforced reasoning phase serve as a pivotal driving force for advancing model reasoning to its limits. Next, we collate existing datasets and evaluation benchmarks currently used for RL fine-tuning, spanning human-annotated datasets, AI-assisted preference data, and program-verification-style corpora. Subsequently, we review the mainstream open-source tools and training frameworks available, providing clear practical references for subsequent research. Finally, we analyse the future challenges and trends in the field of RL-enhanced LLMs. This survey aims to present researchers and practitioners with the latest developments and frontier trends at the intersection of RL and LLMs, with the goal of fostering the evolution of LLMs that are more intelligent, generalizable, and secure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。