用因果推理定位多智能体系统故障根源,准确率超36%。
Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- 基于反向数据流与谢尔平利值,实现多粒度因果归因。
- 在Who&When等基准上达36.2%的步骤级故障定位准确率。
- 适合需可解释性与自动优化的复杂多智能体系统研发者。
多智能体系统(MAS)在自动化复杂任务中至关重要,但其实际部署受故障归因难题严重制约。现有依赖统计相关性的诊断工具效果有限,在Who&When等挑战性基准上,最先进方法的根因步骤定位准确率不足15%。为此,我们提出首个基于多粒度因果推理的故障归因框架。核心贡献包括:(1) 性能因果倒置原则,通过反转执行日志中的数据流,结合谢尔平利值精准分配代理责任;(2) 一种新因果发现算法CDC-MAS,有效应对多智能体交互数据的非平稳性,稳健识别关键失败步骤。归因结果直接驱动自动化优化闭环,生成建议并通过反事实模拟验证有效性。在Who&When和TRAIL基准上的评估显示显著提升:方法最高达36.2%的步骤级准确率,生成的优化使整体任务成功率平均提高22.4%。本工作为调试复杂智能体交互提供了原则性且高效的方法,推动更可靠、可解释的多智能体系统发展。
原文摘要 · Abstract (English)
Multi-agent systems (MAS) are critical for automating complex tasks, yet their practical deployment is severely hampered by the challenge of failure attribution. Current diagnostic tools, which rely on statistical correlations, are fundamentally inadequate; on challenging benchmarks like Who\&When, state-of-the-art methods achieve less than 15\% accuracy in locating the root-cause step of a failure. To address this critical gap, we introduce the first failure attribution framework for MAS grounded in multi-granularity causal inference. Our approach makes two key technical contributions: (1) a performance causal inversion principle, which correctly models performance dependencies by reversing the data flow in execution logs, combined with Shapley values to accurately assign agent-level blame; (2) a novel causal discovery algorithm, CDC-MAS, that robustly identifies critical failure steps by tackling the non-stationary nature of MAS interaction data. The framework's attribution results directly fuel an automated optimization loop, generating targeted suggestions whose efficacy is validated via counterfactual simulations. Evaluations on the Who\&When and TRAIL benchmarks demonstrate a significant leap in performance. Our method achieves up to 36.2\% step-level accuracy. Crucially, the generated optimizations boost overall task success rates by an average of 22.4\%. This work provides a principled and effective solution for debugging complex agent interactions, paving the way for more reliable and interpretable multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。