arXiv:2510.16720cs.AI2025-10综述被引 11

从外挂式流程到模型内生智能,让AI自己规划、用工具、记经验。

Beyond Pipelines: A Survey of the Paradigm Shift toward Model-Native Agentic AI

  • 用强化学习让模型自主探索,不再依赖外部脚本
  • 规划、工具使用、记忆等能力已实现端到端学习
  • 适合研究智能体系统架构与下一代AI模型的学者

代理型AI的快速发展标志着人工智能的新阶段:大语言模型不再仅作回应,而是能行动、推理与适应。本文梳理了构建智能体的范式转变——从依赖外部逻辑编排规划、工具调用和记忆的流水线系统,转向将这些能力内化于模型参数中的新型模型原生范式。首先,以强化学习(RL)作为推动该转变的算法引擎,通过将学习目标从模仿静态数据转向以结果为导向的探索,实现了语言、视觉及具身领域中LLM+RL+任务的统一解决方案。在此基础上,系统回顾了规划、工具使用和记忆三项能力如何由外部脚本模块演变为端到端学习的行为。此外,还分析了这一范式转变对主流应用的影响,如强调长程推理的深度研究智能体与注重具身交互的GUI智能体。最后,探讨了多智能体协作与反思等能力的持续内化,以及未来系统层与模型层角色的演变。整体表明,模型原生代理型AI正朝着集学习与交互于一体的整合框架演进,标志着从构建应用智能的系统,转向让模型通过经验自我成长的范式跃迁。

原文摘要 · Abstract (English)

The rapid evolution of agentic AI marks a new phase in artificial intelligence, where Large Language Models (LLMs) no longer merely respond but act, reason, and adapt. This survey traces the paradigm shift in building agentic AI: from Pipeline-based systems, where planning, tool use, and memory are orchestrated by external logic, to the emerging Model-native paradigm, where these capabilities are internalized within the model's parameters. We first position Reinforcement Learning (RL) as the algorithmic engine enabling this paradigm shift. By reframing learning from imitating static data to outcome-driven exploration, RL underpins a unified solution of LLM + RL + Task across language, vision and embodied domains. Building on this, the survey systematically reviews how each capability -- Planning, Tool use, and Memory -- has evolved from externally scripted modules to end-to-end learned behaviors. Furthermore, it examines how this paradigm shift has reshaped major agent applications, specifically the Deep Research agent emphasizing long-horizon reasoning and the GUI agent emphasizing embodied interaction. We conclude by discussing the continued internalization of agentic capabilities like Multi-agent collaboration and Reflection, alongside the evolving roles of the system and model layers in future agentic AI. Together, these developments outline a coherent trajectory toward model-native agentic AI as an integrated learning and interaction framework, marking the transition from constructing systems that apply intelligence to developing models that grow intelligence through experience.

智能体强化学习模型内生代理架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。