多智能体架构让AI应急响应推荐全部可执行,准确率提升140倍。
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
- 用多个AI智能体协作替代单个AI,提升决策质量
- 多智能体推荐100%可执行,正确率比单智能体高140倍
- 结果稳定无波动,适合生产环境部署
大型语言模型(LLMs)有望加速生产系统中的事件响应,但单智能体方法常生成模糊、不可用的建议。我们提出MyAntFarm.ai,一个可复现的容器化框架,证明多智能体编排能从根本上提升基于LLM的事件响应质量。在348次控制实验中,对比单智能体协作者与多智能体系统在相同事件场景下的表现,发现多智能体编排实现100%可操作建议率,而单智能体仅为1.7%,行动具体性提升80倍,解决方案正确率提升140倍。关键的是,多智能体系统在所有试验中均无质量波动,实现了单智能体无法达成的生产级服务等级协议(SLA)承诺。两架构理解延迟相近(约40秒),表明其价值在于确定性质量而非速度。我们引入决策质量(DQ)这一新指标,捕捉有效性、具体性和正确性等运维部署必需属性,现有LLM评估指标未涵盖此维度。研究结果将多智能体编排从性能优化重构为基于LLM的事件响应的生产就绪必要条件。所有代码、Docker配置及试验数据均公开可复现。
原文摘要 · Abstract (English)
Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable recommendations. We present MyAntFarm.ai, a reproducible containerized framework demonstrating that multi-agent orchestration fundamentally transforms LLM-based incident response quality. Through 348 controlled trials comparing single-agent copilot versus multi-agent systems on identical incident scenarios, we find that multi-agent orchestration achieves 100% actionable recommendation rate versus 1.7% for single-agent approaches, an 80 times improvement in action specificity and 140 times improvement in solution correctness. Critically, multi-agent systems exhibit zero quality variance across all trials, enabling production SLA commitments impossible with inconsistent single-agent outputs. Both architectures achieve similar comprehension latency (approx.40s), establishing that the architectural value lies in deterministic quality, not speed. We introduce Decision Quality (DQ), a novel metric capturing validity, specificity, and correctness properties essential for operational deployment that existing LLM metrics do not address. These findings reframe multi-agent orchestration from a performance optimization to a production-readiness requirement for LLM-based incident response. All code, Docker configurations, and trial data are publicly available for reproduction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。