arXiv:2511.15755cs.AIcs.SE2025-11被引 8

多智能体架构让AI应急响应推荐全部可执行,准确率提升140倍。

Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response

  • 用多个AI智能体协作替代单个AI,提升决策质量
  • 多智能体推荐100%可执行,正确率比单智能体高140倍
  • 结果稳定无波动,适合生产环境部署

大型语言模型(LLMs)有望加速生产系统中的事件响应,但单智能体方法常生成模糊、不可用的建议。我们提出MyAntFarm.ai,一个可复现的容器化框架,证明多智能体编排能从根本上提升基于LLM的事件响应质量。在348次控制实验中,对比单智能体协作者与多智能体系统在相同事件场景下的表现,发现多智能体编排实现100%可操作建议率,而单智能体仅为1.7%,行动具体性提升80倍,解决方案正确率提升140倍。关键的是,多智能体系统在所有试验中均无质量波动,实现了单智能体无法达成的生产级服务等级协议(SLA)承诺。两架构理解延迟相近(约40秒),表明其价值在于确定性质量而非速度。我们引入决策质量(DQ)这一新指标,捕捉有效性、具体性和正确性等运维部署必需属性,现有LLM评估指标未涵盖此维度。研究结果将多智能体编排从性能优化重构为基于LLM的事件响应的生产就绪必要条件。所有代码、Docker配置及试验数据均公开可复现。

原文摘要 · Abstract (English)

Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable recommendations. We present MyAntFarm.ai, a reproducible containerized framework demonstrating that multi-agent orchestration fundamentally transforms LLM-based incident response quality. Through 348 controlled trials comparing single-agent copilot versus multi-agent systems on identical incident scenarios, we find that multi-agent orchestration achieves 100% actionable recommendation rate versus 1.7% for single-agent approaches, an 80 times improvement in action specificity and 140 times improvement in solution correctness. Critically, multi-agent systems exhibit zero quality variance across all trials, enabling production SLA commitments impossible with inconsistent single-agent outputs. Both architectures achieve similar comprehension latency (approx.40s), establishing that the architectural value lies in deterministic quality, not speed. We introduce Decision Quality (DQ), a novel metric capturing validity, specificity, and correctness properties essential for operational deployment that existing LLM metrics do not address. These findings reframe multi-agent orchestration from a performance optimization to a production-readiness requirement for LLM-based incident response. All code, Docker configurations, and trial data are publicly available for reproduction.

多智能体应急响应LLM应用确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。