arXiv:2510.13561cs.SEcs.AI2025-10

为SRE设计的AI协作框架,能自动诊断复杂系统故障

OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies

  • 专为SRE诊断设计的多智能体协同架构,支持深度因果推理
  • 在蚂蚁集团日均服务3000人,故障诊断准确率显著优于现有方案
  • 开源可扩展,适合需要自动化运维的大型科技企业

现代软件系统日益复杂,给站点可靠性工程(SRE)团队带来不可持续的运维负担,亟需能模拟专家诊断思维的AI自动化能力。现有解决方案,包括传统AI方法和通用多智能体系统,或缺乏深层因果推理能力,或不适用于SRE特有的调查工作流。为此,我们提出OpenDerisk,一个专为SRE设计的开源多智能体框架。该框架集成诊断导向的协作模型、可插拔推理引擎、知识引擎及标准化协议(MCP),使专业智能体能协同解决跨领域复杂问题。全面评估显示,OpenDerisk在准确率和效率上均显著优于当前最优基线。其工业级可扩展性已在蚂蚁集团得到验证:服务超过3000名日常用户,覆盖多样化场景,证明了实际应用价值。项目已开源,地址为https://github.com/derisk-ai/OpenDerisk/

原文摘要 · Abstract (English)

The escalating complexity of modern software imposes an unsustainable operational burden on Site Reliability Engineering (SRE) teams, demanding AI-driven automation that can emulate expert diagnostic reasoning. Existing solutions, from traditional AI methods to general-purpose multi-agent systems, fall short: they either lack deep causal reasoning or are not tailored for the specialized, investigative workflows unique to SRE. To address this gap, we present OpenDerisk, a specialized, open-source multi-agent framework architected for SRE. OpenDerisk integrates a diagnostic-native collaboration model, a pluggable reasoning engine, a knowledge engine, and a standardized protocol (MCP) to enable specialist agents to collectively solve complex, multi-domain problems. Our comprehensive evaluation demonstrates that OpenDerisk significantly outperforms state-of-the-art baselines in both accuracy and efficiency. This effectiveness is validated by its large-scale production deployment at Ant Group, where it serves over 3,000 daily users across diverse scenarios, confirming its industrial-grade scalability and practical impact. OpenDerisk is open source and available at https://github.com/derisk-ai/OpenDerisk/

SRE多智能体故障诊断开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。