arXiv:2602.09937cs.AIcs.DC2026-02被引 6

剖析大模型在云系统故障诊断中的失败原因,发现共性缺陷

Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?

  • 通过1675次实验分析大模型代理的推理、通信与环境交互问题
  • 发现幻觉数据解读和探索不全是普遍存在的核心错误,跨模型普遍存在
  • 提示工程无效,改进通信协议可降低15%的沟通失败率

大规模云系统故障导致巨大经济损失,自动化根因分析(RCA)对运维稳定至关重要。近期研究利用大语言模型(LLM)代理实现该任务,但现有系统检测准确率仍低,且当前评估框架仅关注最终答案正确性,无法揭示推理失败原因。本文对基于LLM的RCA代理进行过程级失败分析,在五个LLM模型上运行完整OpenRCA基准测试,共产生1,675次代理运行,并将观察到的失败归类为12种陷阱类型,涵盖代理内部推理、代理间通信及代理-环境交互。分析显示,最普遍的陷阱,如幻觉数据解读和不完整探索,无论模型能力高低均普遍存在,表明这些失败源于共享的代理架构而非单个模型局限。受控缓解实验进一步表明,仅靠提示工程无法解决主导性陷阱,而优化代理间通信协议可使通信相关失败降低最多达15个百分点。本研究提出的陷阱分类体系与诊断方法,为设计更可靠的云RCA自主代理提供了基础。

原文摘要 · Abstract (English)

Failures in large-scale cloud systems incur substantial financial losses, making automated Root Cause Analysis (RCA) essential for operational stability. Recent efforts leverage Large Language Model (LLM) agents to automate this task, yet existing systems exhibit low detection accuracy even with capable models, and current evaluation frameworks assess only final answer correctness without revealing why the agent's reasoning failed. This paper presents a process level failure analysis of LLM-based RCA agents. We execute the full OpenRCA benchmark across five LLM models, producing 1,675 agent runs, and classify observed failures into 12 pitfall types across intra-agent reasoning, inter-agent communication, and agent-environment interaction. Our analysis reveals that the most prevalent pitfalls, notably hallucinated data interpretation and incomplete exploration, persist across all models regardless of capability tier, indicating that these failures originate from the shared agent architecture rather than from individual model limitations. Controlled mitigation experiments further show that prompt engineering alone cannot resolve the dominant pitfalls, whereas enriching the inter-agent communication protocol reduces communication-related failures by up to 15 percentage points. The pitfall taxonomy and diagnostic methodology developed in this work provide a foundation for designing more reliable autonomous agents for cloud RCA.

故障诊断大模型代理LLM应用系统可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。