用系统内代理自检,高效发现多智能体系统故障。
POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems

- 让系统内代理互相提问诊断,利用认知多样性自检。
- 复杂任务下检测准确率提升60%(OR=1.60,p=0.008)。
- 适合安全关键场景的多智能体系统开发者使用。
将大语言模型集成到多智能体系统(LLM-MAS)中虽带来强大推理能力,但涌现的故障与幻觉难以界定,阻碍其在安全关键领域的应用——这一问题在新兴人工智能法规下愈发不可接受。现有评估方法存在共性缺陷:集中式判断易形成单点故障,且需领域专业知识。本文提出POIROT协议,将系统自身代理作为诊断层,利用架构中已有的认知多样性。在多种测试场景中,POIROT超越单一语言模型评估基线,性能随问题复杂度、代理数量和故障维度提升,在复合故障条件下仍有效。结果表明,安全监控无需外部依赖:执行角色的代理具备足够的集体智能完成自我审计。我们开源了POIROT,并配套发布BLAME基准,用于安全关键多智能体系统的故障归因评估。
原文摘要 · Abstract (English)
Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation. Existing evaluation paradigms share a common flaw: centralised judgment creates single points of failure and demands domain-specific expertise. Here we present POIROT, a protocol that repurposes a system's own agents as its diagnostic layer, leveraging the epistemic diversity already present in the architecture. Across evaluated settings, POIROT outperforms single-LLM evaluator baselines, with gains that scale with problem complexity (OR = 1.60, $p = 0.008$), agent count, and fault dimensionality, persisting under compound fault conditions. These results demonstrate that safety oversight need not be externalised: the agents executing a role carry sufficient collective intelligence to audit it. We release POIROT as an open-source library alongside BLAME, a benchmark for fault attribution in safety-critical multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。