系统梳理智能体适应性研究,涵盖训练后优化与技能记忆体系。
Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
- 提出四范式框架,分代理与工具双路径改进智能体能力。
- 验证可验证奖励下强化学习显著提升推理与工具使用性能。
- 适合关注智能体持续学习与安全部署的研究者与开发者。
大型语言模型(LLM)智能体正超越单纯提示阶段。ChatGPT标志着通用型LLM助手的兴起,DeepSeek证明了基于可验证奖励的策略强化学习能提升推理与工具使用能力,OpenClaw则展示了智能体积累持久记忆与可复用技能的新方向。然而,后训练、检索、记忆与技能系统的研究仍分散割裂。本综述以‘适应’为核心概念,统一分析代理及其工具在预训练后的优化机制。构建四范式框架:代理侧包含A1(工具执行信号驱动)与A2(代理输出信号驱动),通过监督微调、偏好优化与强化学习改进自身;工具侧包含T1(代理无关)与T2(代理监督),前者提供通用预训练模块,后者利用代理输出训练记忆系统、技能库或轻量级子代理。基于该框架,我们回顾后训练方法、自适应记忆架构与代理技能,比较其在成本、灵活性与泛化性上的权衡,并总结深度研究、软件开发、计算机使用及药物发现等领域的评估实践。最后,指出代理-工具协同适应、持续学习、安全性与高效部署中的开放问题。
原文摘要 · Abstract (English)
Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。