arXiv:2509.13782cs.SEcs.AI2025-09被引 34

通过频谱分析自动定位多智能体系统失败根源

Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis

  • 基于多次执行轨迹的变异分析,量化每个动作的可疑度
  • 在Who and When基准上优于12种基线方法,准确率显著提升
  • 适合需要高效调试复杂多智能体系统的研究人员

大型语言模型驱动的多智能体系统(MASs)正被广泛应用于编程和科学发现等复杂任务。然而,其失败归因——即识别导致失败的具体智能体行为——仍缺乏系统方法且耗时费力。为此,我们提出FAMAS,首个基于频谱分析的多智能体系统失败归因方法。该方法通过系统性轨迹重播与抽象,结合频谱分析,估计各智能体动作在多次执行中引发失败的可能性。核心创新在于设计了一种专为多智能体系统定制的可疑度公式,融合智能体行为模式与动作行为模式,捕捉执行轨迹中的激活特征。在Who and When基准上,通过与12种基线方法的对比实验,FAMAS展现出卓越性能,全面超越现有方法。

原文摘要 · Abstract (English)

Large Language Model Powered Multi-Agent Systems (MASs) are increasingly employed to automate complex real-world problems, such as programming and scientific discovery. Despite their promising, MASs are not without their flaws. However, failure attribution in MASs - pinpointing the specific agent actions responsible for failures - remains underexplored and labor-intensive, posing significant challenges for debugging and system improvement. To bridge this gap, we propose FAMAS, the first spectrum-based failure attribution approach for MASs, which operates through systematic trajectory replay and abstraction, followed by spectrum analysis.The core idea of FAMAS is to estimate, from variations across repeated MAS executions, the likelihood that each agent action is responsible for the failure. In particular, we propose a novel suspiciousness formula tailored to MASs, which integrates two key factor groups, namely the agent behavior group and the action behavior group, to account for the agent activation patterns and the action activation patterns within the execution trajectories of MASs. Through expensive evaluations against 12 baselines on the Who and When benchmark, FAMAS demonstrates superior performance by outperforming all the methods in comparison.

多智能体失败归因频谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。