arXiv:2509.02547cs.AIcs.CL2025-09综述被引 197

将大模型从文本生成器变为能自主决策的智能体,探索其强化学习新范式。

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

  • 提出双维度分类体系,涵盖规划、工具使用等核心能力
  • 揭示强化学习如何让智能体行为从静态模块升级为自适应系统
  • 整合500+最新研究,提供开源环境与评测基准合集

智能体强化学习(Agentic RL)的兴起标志着大语言模型强化学习(LLM RL)范式的根本转变,将大模型从被动的序列生成器重塑为嵌入复杂动态世界中的自主决策智能体。本文通过对比传统单步马尔可夫决策过程(MDP)与更复杂的时序扩展、部分可观测马尔可夫决策过程(POMDP),正式确立这一概念转型。在此基础上,提出双重分类框架:一以核心智能体能力(如规划、工具使用、记忆、推理、自我改进、感知)为维度,二以跨任务领域的应用为维度。核心论点是,强化学习是实现这些能力从静态启发式模块向自适应、鲁棒智能体行为转化的关键机制。为推动未来研究,本综述系统整理了开源环境、评测基准与框架资源。基于对超过500篇近期文献的综合分析,勾勒出该快速演进领域的发展轮廓,并指明塑造可扩展通用人工智能智能体的机遇与挑战。

原文摘要 · Abstract (English)

The emergence of agentic reinforcement learning (Agentic RL) marks a paradigm shift from conventional reinforcement learning applied to large language models (LLM RL), reframing LLMs from passive sequence generators into autonomous, decision-making agents embedded in complex, dynamic worlds. This survey formalizes this conceptual shift by contrasting the degenerate single-step Markov Decision Processes (MDPs) of LLM-RL with the temporally extended, partially observable Markov decision processes (POMDPs) that define Agentic RL. Building on this foundation, we propose a comprehensive twofold taxonomy: one organized around core agentic capabilities, including planning, tool use, memory, reasoning, self-improvement, and perception, and the other around their applications across diverse task domains. Central to our thesis is that reinforcement learning serves as the critical mechanism for transforming these capabilities from static, heuristic modules into adaptive, robust agentic behavior. To support and accelerate future research, we consolidate the landscape of open-source environments, benchmarks, and frameworks into a practical compendium. By synthesizing over five hundred recent works, this survey charts the contours of this rapidly evolving field and highlights the opportunities and challenges that will shape the development of scalable, general-purpose AI agents.

智能体强化学习大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。