梳理大模型智能体在工具使用、规划与推理中的六大失败模式。
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
- 从27篇论文中归纳出智能体推理-行动链各阶段的失效模式
- 任务越长,失败越严重,子任务表现好不等于整体成功
- 适合研究者和开发者参考,避免重复踩坑
大语言模型智能体在工具使用、多步规划、多智能体协作及长时程任务中的评估日益增多,但报告的性能提升常掩盖了跨不同评测中反复出现的失败模式。本文综合2023至2026年的27篇基准测试、分类体系与审计论文,覆盖19个不同基准,构建了一个涵盖工具使用、规划、长时程推理、多智能体协作、安全性和测量有效性六个维度的统一失效分类体系。我们识别出六类失败集群:(1) 工具调用与参数错误,(2) 规划与约束满足失败,(3) 上下文累积导致的长时程退化,(4) 多智能体协作失败,(5) 在对抗或条件不明确下的安全与隐私问题,(6) 评估有效性缺陷。该分类通过迭代归类独立报告的错误类型,对应智能体推理-行动流程中的不同阶段。研究发现,失败随任务长度呈非线性叠加,单个子任务表现优异无法保证端到端成功,额外的结构支持也未显著提升可靠性。同时,在单轮工具调用、短时程网页导航和特定编码任务中已取得显著进展。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts. This paper synthesizes 27 benchmark, taxonomy, and audit papers (2023-2026), spanning 19 distinct benchmarks, into a cross-cutting taxonomy of agent limitations. To our knowledge, this is the first synthesis that integrates evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and measurement validity into a single, unified taxonomy of LLM agent limitations. We identify six failure clusters: (1) tool invocation and parameter-level errors, (2) planning and constraint-satisfaction failures, (3) long-horizon degradation from context accumulation, (4) multi-agent coordination failures, (5) safety and security failures under adversarial or underspecified conditions, and (6) measurement validity problems. The taxonomy was derived iteratively by grouping independently reported error categories into themes corresponding to distinct stages of the agent reasoning-to-action pipeline. Across the literature, we find that failures compound nonlinearly with task length, that strong performance on individual sub-tasks does not reliably translate into end-to-end success, and that additional scaffolding does not consistently improve reliability. At the same time, substantial progress has been demonstrated in single-turn tool use, short-horizon web navigation, and narrowly scoped coding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。