系统梳理智能体架构与评估挑战,助力自然语言对接真实计算。
AI Agent Systems: Architectures, Applications, and Evaluation
- 构建包含推理、规划、工具调用的统一智能体框架
- 揭示延迟与准确率、自主性与可控性间的权衡关系
- 适合研究智能体系统设计与评估的开发者与学者
AI智能体——结合基础模型与推理、规划、记忆及工具使用的系统——正迅速成为自然语言意图与真实世界计算之间的实用接口。本文综述了智能体架构的新兴领域,涵盖:(i) 深思与推理(如思维链分解、自我反思与验证、约束感知决策),(ii) 规划与控制(从反应式策略到分层与多步规划器),以及 (iii) 工具调用与环境交互(检索、代码执行、API 调用、多模态感知)。我们将已有工作归纳为统一分类体系,涵盖智能体组件(策略/大模型核心、记忆、世界模型、规划器、工具路由、评议员)、编排模式(单智能体与多智能体;集中式与分布式协调)以及部署场景(离线分析与在线交互辅助;安全关键与开放任务)。讨论关键设计权衡——延迟与精度、自主性与可控性、能力与可靠性——并指出评估复杂性源于非确定性、长程信用分配、工具与环境变异性,以及重试和上下文增长等隐含成本。最后总结测量与基准实践(任务套件、人类偏好与效用指标、约束下成功率、鲁棒性与安全性),并提出开放挑战,包括工具动作的验证与防护机制、可扩展的记忆与上下文管理、智能体决策的可解释性,以及在真实负载下的可复现评估。
原文摘要 · Abstract (English)
AI agents -- systems that combine foundation models with reasoning, planning, memory, and tool use -- are rapidly becoming a practical interface between natural-language intent and real-world computation. This survey synthesizes the emerging landscape of AI agent architectures across: (i) deliberation and reasoning (e.g., chain-of-thought-style decomposition, self-reflection and verification, and constraint-aware decision making), (ii) planning and control (from reactive policies to hierarchical and multi-step planners), and (iii) tool calling and environment interaction (retrieval, code execution, APIs, and multimodal perception). We organize prior work into a unified taxonomy spanning agent components (policy/LLM core, memory, world models, planners, tool routers, and critics), orchestration patterns (single-agent vs.\ multi-agent; centralized vs.\ decentralized coordination), and deployment settings (offline analysis vs.\ online interactive assistance; safety-critical vs.\ open-ended tasks). We discuss key design trade-offs -- latency vs.\ accuracy, autonomy vs.\ controllability, and capability vs.\ reliability -- and highlight how evaluation is complicated by non-determinism, long-horizon credit assignment, tool and environment variability, and hidden costs such as retries and context growth. Finally, we summarize measurement and benchmarking practices (task suites, human preference and utility metrics, success under constraints, robustness and security) and identify open challenges including verification and guardrails for tool actions, scalable memory and context management, interpretability of agent decisions, and reproducible evaluation under realistic workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。