系统梳理大模型智能体泛化能力,指明提升路径与研究方向
Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- 构建层级任务-领域框架,明确定义智能体泛化边界
- 归纳三类提升泛化的方法:模型、组件及交互优化
- 提出可泛化框架与智能体的转化路径,适合研究者参考
基于大语言模型(LLM)的智能体已成为超越文本生成的新范式,通过融合推理、感知、记忆与工具使用,在网页导航、家庭机器人等多领域部署。其核心挑战在于泛化能力——在未见过的任务、环境与领域中保持稳定表现,尤其超出训练数据范围时。尽管关注度上升,当前对智能体泛化仍缺乏明确定义,也无系统评估与提升方法。本文首次全面综述该问题,首先强调其重要性并基于层级任务-领域本体界定边界;接着梳理数据集、评估维度与指标,指出其局限;然后将提升方法分为三类:基础模型、智能体组件及其交互;进一步区分可泛化框架与可泛化智能体,并阐明前者如何转化为后者;最后提出关键挑战与未来方向,包括建立标准化框架、方差与成本导向指标,以及方法论创新与架构设计融合。本综述旨在为可靠跨应用智能体研发奠定理论基础。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based agents have emerged as a new paradigm that extends LLMs' capabilities beyond text generation to dynamic interaction with external environments. By integrating reasoning with perception, memory, and tool use, agents are increasingly deployed in diverse domains like web navigation and household robotics. A critical challenge, however, lies in ensuring agent generalizability - the ability to maintain consistent performance across varied instructions, tasks, environments, and domains, especially those beyond agents' fine-tuning data. Despite growing interest, the concept of generalizability in LLM-based agents remains underdefined, and systematic approaches to measure and improve it are lacking. In this survey, we provide the first comprehensive review of generalizability in LLM-based agents. We begin by emphasizing agent generalizability's importance by appealing to stakeholders and clarifying the boundaries of agent generalizability by situating it within a hierarchical domain-task ontology. We then review datasets, evaluation dimensions, and metrics, highlighting their limitations. Next, we categorize methods for improving generalizability into three groups: methods for the backbone LLM, for agent components, and for their interactions. Moreover, we introduce the distinction between generalizable frameworks and generalizable agents and outline how generalizable frameworks can be translated into agent-level generalizability. Finally, we identify critical challenges and future directions, including developing standardized frameworks, variance- and cost-based metrics, and approaches that integrate methodological innovations with architecture-level designs. By synthesizing progress and highlighting opportunities, this survey aims to establish a foundation for principled research on building LLM-based agents that generalize reliably across diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。