arXiv:2603.25001cs.AI2026-03被引 11

提出多视角故障归因新范式,打破单一原因假设。

Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation

  • 引入多视角故障归因,考虑复杂系统中多个可能原因。
  • 构建首个支持多视角的基准测试MP-Bench及对应评估协议。
  • 揭示旧结论误判源于基准设计缺陷,强调真实调试需多视角评估。

故障归因对诊断和改进多智能体系统(MAS)至关重要,但现有基准和方法大多假设每个故障仅有一个确定的根本原因。实际上,由于智能体间复杂的依赖关系和模糊的执行轨迹,MAS故障常存在多个合理归因。本文从多视角出发重新审视故障归因问题,提出多视角故障归因这一新范式,明确考虑归因的不确定性。为此,我们引入MP-Bench——首个面向多视角故障归因的基准数据集,并设计了适配该范式的新型评估协议。通过大量实验发现,以往认为大语言模型在故障归因上表现不佳的结论,主要源于现有基准设计的局限性。结果表明,为实现真实可靠的多智能体系统调试,必须采用多视角基准与评估方法。

原文摘要 · Abstract (English)

Failure attribution is essential for diagnosing and improving multi-agent systems (MAS), yet existing benchmarks and methods largely assume a single deterministic root cause for each failure. In practice, MAS failures often admit multiple plausible attributions due to complex inter-agent dependencies and ambiguous execution trajectories. We revisit MAS failure attribution from a multi-perspective standpoint and propose multi-perspective failure attribution, a practical paradigm that explicitly accounts for attribution ambiguity. To support this setting, we introduce MP-Bench, the first benchmark designed for multi-perspective failure attribution in MAS, along with a new evaluation protocol tailored to this paradigm. Through extensive experiments, we find that prior conclusions suggesting LLMs struggle with failure attribution are largely driven by limitations in existing benchmark designs. Our results highlight the necessity of multi-perspective benchmarks and evaluation protocols for realistic and reliable MAS debugging.

多智能体故障归因评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。