arXiv:2508.13143cs.AIcs.SE2025-08中稿 · ASE 2025 NIER被引 19

分析智能体任务失败原因,提出可落地的改进方案。

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

  • 构建34个编程任务基准,系统评估智能体表现。
  • 仅约50%任务成功,主要败在规划错误与响应生成失误。
  • 给出分阶段失败归因框架,适合研究者优化智能体设计。

由大语言模型驱动的自主智能体系统在自动化复杂任务方面展现出巨大潜力。然而,当前评估多依赖成功率,缺乏对系统内交互、通信机制及失败原因的系统性分析。为此,我们设计了一个包含34个代表性可编程任务的基准,用于严格评估自主智能体。基于该基准,我们评估了三种主流开源智能体框架搭配两种LLM骨干网络,观察到任务完成率约为50%。通过深入的失败分析,我们提出了一个与任务阶段对应的三层失败分类体系,揭示出规划错误、任务执行问题和响应生成错误是主要瓶颈。基于此,我们提出了提升智能体规划与自诊断能力的具体改进策略。本研究的失败分类框架与缓解建议,为未来构建更鲁棒、高效的自主智能体系统提供了实证基础。

原文摘要 · Abstract (English)

Autonomous agent systems powered by Large Language Models (LLMs) have demonstrated promising capabilities in automating complex tasks. However, current evaluations largely rely on success rates without systematically analyzing the interactions, communication mechanisms, and failure causes within these systems. To bridge this gap, we present a benchmark of 34 representative programmable tasks designed to rigorously assess autonomous agents. Using this benchmark, we evaluate three popular open-source agent frameworks combined with two LLM backbones, observing a task completion rate of approximately 50%. Through in-depth failure analysis, we develop a three-tier taxonomy of failure causes aligned with task phases, highlighting planning errors, task execution issues, and incorrect response generation. Based on these insights, we propose actionable improvements to enhance agent planning and self-diagnosis capabilities. Our failure taxonomy, together with mitigation advice, provides an empirical foundation for developing more robust and effective autonomous agent systems in the future.

智能体失败分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。