arXiv:2412.10400cs.CLcs.AI2024-12综述被引 89

系统梳理强化学习增强大模型的技术演进与挑战

Reinforcement Learning Enhanced LLMs: A Survey

  • 从基础原理到主流方法,全面解析RL增强LLM的技术路径
  • 对比分析RLHF、RLAIF与DPO等关键对齐技术的优劣
  • 适合希望快速掌握该领域前沿进展的研究者阅读

强化学习(RL)增强的大语言模型(LLMs),以DeepSeek-R1为代表,展现出卓越性能。尽管在提升模型能力方面有效,其实现过程仍极为复杂,涉及复杂的算法、奖励建模策略与优化技术,给研究者和实践者带来系统理解上的困难。当前该领域缺乏全面综述,制约了进一步发展。本文系统梳理最新的研究成果,涵盖RL基础、主流增强模型、基于奖励模型的RLHF与RLAIF技术,以及不依赖奖励模型的直接偏好优化(DPO)方法。同时指出现有方法的不足,提出未来改进方向。项目主页见:https://github.com/ShuheWang1998/Reinforcement-Learning-Enhanced-LLMs-A-Survey。

原文摘要 · Abstract (English)

Reinforcement learning (RL) enhanced large language models (LLMs), particularly exemplified by DeepSeek-R1, have exhibited outstanding performance. Despite the effectiveness in improving LLM capabilities, its implementation remains highly complex, requiring complex algorithms, reward modeling strategies, and optimization techniques. This complexity poses challenges for researchers and practitioners in developing a systematic understanding of RL-enhanced LLMs. Moreover, the absence of a comprehensive survey summarizing existing research on RL-enhanced LLMs has limited progress in this domain, hindering further advancements. In this work, we are going to make a systematic review of the most up-to-date state of knowledge on RL-enhanced LLMs, attempting to consolidate and analyze the rapidly growing research in this field, helping researchers understand the current challenges and advancements. Specifically, we (1) detail the basics of RL; (2) introduce popular RL-enhanced LLMs; (3) review researches on two widely-used reward model-based RL techniques: Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF); and (4) explore Direct Preference Optimization (DPO), a set of methods that bypass the reward model to directly use human preference data for aligning LLM outputs with human expectations. We will also point out current challenges and deficiencies of existing methods and suggest some avenues for further improvements. Project page of this work can be found at https://github.com/ShuheWang1998/Reinforcement-Learning-Enhanced-LLMs-A-Survey.

强化学习大模型对齐RLHFDPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。