arXiv:2608.05263cs.AI2026-08

评测多智能体编排的失败模式与恢复能力,揭示关键故障点。

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

  • 通过可控故障注入测试编排系统在企业工作流中的表现
  • 关键词路由在对抗性场景下准确率0%,意图推理模型达100%
  • 发现三类故障处理层级,隐式语义错误无法恢复

多智能体编排框架正从演示走向生产,但现有基准通常只报告任务准确率,缺乏对失败原因、级联起点及路由决策失误的诊断。OrchestraBench 通过受控、可复现的故障注入机制,在模板化的企业工作流上评估失败、恢复与分解质量。引入级联半径和按故障模式的恢复率作为核心指标,采用置信区间与配对检验比较路由策略。在26个带标签的诊断案例中,关键词/标志路由在误导或缺失表面标志的对抗性情况下准确率为0%,而意图推理模型路由达到100%,与理想值一致。在真实Claude智能体上的可验证算术依赖链机制探针显示,五种MAST模式存在三级故障处理能力:工具故障完全恢复(1.0),模糊委派部分恢复(0.30),三种隐式或语义模式均未恢复(0.0)。该顺序在贷款审批工作流重构及Sonnet、Opus、Haiku模型间保持一致,尽管绝对成功率随上下文变化。盲目重试会重现隐式故障并延迟检测,表明检测与归因对遏制至关重要。级联半径随流水线深度增加(深度3-7时均值从0.9升至4.7)。可信状态修复消融实验表明,看似有效的遏制效果主要源于可信状态信号,而非自主检测。这些结果基于受控链机制探针,非领域或负载声明。

原文摘要 · Abstract (English)

Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench evaluates failure, recovery, and decomposition through a controlled, seed-reproducible failure-injection harness over templated enterprise workflows. It introduces cascade radius and per-failure-mode recovery as primary metrics and compares routing policies with bootstrap confidence intervals and paired tests. On a 26-case gold-labelled diagnostic, a keyword/flag router scored 0% on adversarial cases with misleading or missing surface flags, whereas an intent-reasoning model router scored 100%, matching the oracle. Controlled mechanism probes with a real Claude agent over a verifiable arithmetic dependency chain revealed three failure-handling tiers across five MAST modes: tool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0). This ordering persisted when the computation was reframed as a loan-approval workflow and across Sonnet, Opus, and Haiku, although absolute rates shifted with context. Blind retry reproduced latent faults and increased time to detection, indicating that detection and attribution are necessary for containment. Cascade radius increased with pipeline depth (mean 0.9 to 4.7 across depths 3-7). A trusted-state repair ablation showed that apparent containment gains primarily came from the trusted-state signal rather than autonomous detection. These results are controlled-chain mechanism probes, not domain-workload claims.

多智能体故障诊断编排评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。