arXiv:2604.17312cs.LGcs.AI2026-04ACL综述

系统梳理数据稀缺下大模型强化学习的挑战与解法

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions

论文配图:A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions
图 1 · 摘自论文原文
  • 构建数据、训练、框架三视角融合的分层分析框架
  • 首次系统归纳数据高效强化学习方法并建立分类体系
  • 适合关注大模型后训练效率的研究者和实践者

强化学习(RL)已成为提升大语言模型(LLM)推理能力的重要后训练范式。然而,针对LLM的强化学习面临严重数据稀缺问题,包括高质量外部监督数据有限以及模型生成经验总量受限。这一限制使得数据高效强化学习成为关键研究方向。本文首次系统性地综述了数据稀缺环境下大语言模型的强化学习。我们提出一个自下而上的分层框架,涵盖数据中心、训练中心和框架中心三个互补视角,构建现有方法的分类体系,总结代表性方法并分析其优劣。该分类旨在为数据高效强化学习的设计空间提供清晰的概念基础,并指导该新兴领域研究人员。我们希望本综述能为未来研究提供全面路线图,激发更高效、可扩展的强化学习后训练新方向。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as a powerful post-training paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, reinforcement learning for LLMs faces substantial data scarcity challenges, including the limited availability of high-quality external supervision and the constrained volume of model-generated experience. These limitations make data-efficient reinforcement learning a critical research direction. In this survey, we present the first systematic review of reinforcement learning for LLMs under data scarcity. We propose a bottom-up hierarchical framework built around three complementary perspectives: the data-centric perspective, the training-centric perspective, and the framework-centric perspective. We develop a taxonomy of existing methods, summarize representative approaches in each category, and analyze their strengths and limitations. Our taxonomy aims to provide a clear conceptual foundation for understanding the design space of data-efficient RL for LLMs and to guide researchers working in this emerging area. We hope this survey offers a comprehensive roadmap for future research and inspires new directions toward more efficient and scalable reinforcement learning post-training for LLMs.

强化学习大模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。