arXiv:2607.09996cs.AIcs.MA2026-07被引 3

构建大规模基准,测试大模型能否准确定位智能体失败原因。

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

论文配图:Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
图 1 · 摘自论文原文
  • 通过严格控制流程生成12326条带标签的失败轨迹。
  • 发现不同模态和模型家族在归因上存在系统性差异。
  • 适合研究智能体可解释性与自动化故障诊断的学者。

自动化故障归因利用大语言模型识别智能体系统失败的位置与原因。随着智能体能力提升,其失败表现愈发隐蔽,自动化归因的重要性日益凸显。本文提出Who&When Pro,一个面向智能体系统自动化故障归因的大规模基准。通过严格控制的流水线,在精确重放成功前缀后注入故障,构建了涵盖3种模态、26个基准的12,326条失败轨迹,并配有黄金标注。除基准构建外,我们进行了广泛的实验与分析,揭示了模型在不同模态、协议和模型家族间进行故障归因时的系统性模式,为未来自动化故障归因系统提供了实证指导。

原文摘要 · Abstract (English)

Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmark for automated failure attribution in agentic systems. Using a strictly controlled pipeline that injects a failure only after exactly replaying a successful prefix, we construct 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks covering various scenarios. Beyond benchmarking, we conduct extensive experiments and analyses, revealing systematic patterns in how models attribute failures across modalities, protocols, and model families, and providing empirical guidance for future automated failure attribution systems.

智能体故障归因大模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。