arXiv:2606.09122cs.SEcs.AI2026-06被引 3

用智能代理系统自动处理大规模网络故障,解决人工响应慢的问题。

Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations

论文配图:Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations
图 1 · 摘自论文原文
  • 多智能体协作架构,分工检测、诊断、修复网络故障
  • 90%以上常见故障可自主解决,且有安全回滚机制保障
  • 适合超大规模云运维团队,提升故障响应效率

超大规模云网络基础设施面临故障数量多、速度快、复杂度高的运营挑战,传统人工响应已难以应对。本文提出一种用于大规模网络运维的智能体式人工智能架构,通过多智能体协同,实现故障的自主检测、诊断与修复。系统采用分层智能体分解、基于技能的标准化工具调用、运行手册结构化知识编码、渐进式自治与安全边界控制、闭环验证等设计原则。该架构已在某大型云服务商生产环境部署,实测表明对常见故障类别,自主解决率超过90%,并通过分层授权与回滚机制确保安全性。文中还探讨了设计权衡、失效模式及规模化运行的经验教训。

原文摘要 · Abstract (English)

Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace with the volume, velocity, and complexity of failures. This paper presents an agentic AI architecture for autonomous incident resolution in large-scale network operations. Our system employs a multi-agent orchestration framework where specialized AI agents collaborate to detect, diagnose, and remediate network incidents without human intervention. We describe the architectural principles, including hierarchical agent decomposition, skills-based tool invocation via standardized protocols, structured knowledge encoding from operational runbooks, progressive autonomy with safety boundaries, and closed-loop verification. The architecture has been deployed in production at a major cloud provider, demonstrating that agentic AI systems can achieve autonomous resolution rates exceeding 90% for common incident categories while maintaining safety guarantees through layered authorization and rollback mechanisms. We discuss design tradeoffs, failure modes, and lessons learned from operating autonomous AI agents at scale.

智能体系统网络运维自动化云平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。