arXiv:2606.13543cs.NIcs.LG2026-06

用反事实模拟分析网络故障根源,提升运维决策准确性。

NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks

论文配图:NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks
图 1 · 摘自论文原文
  • 基于图时序建模与反事实推理,自动推断故障传播路径。
  • 在31个真实案例中比规则基线提升16.1%定位准确率。
  • 适合云服务商运维团队快速诊断复杂网络故障。

现有根因分析方法多依赖静态规则、相关性启发或拓扑局部推理,在动态环境中难以泛化。本文提出NetCause,一种自监督学习框架,将网络事件建模为图-时序过程,通过反事实模拟对候选根因进行可解释排序,并可自然集成运维操作。模型在某头部云厂商生产网络中超过1500个故障事件上训练,评估在31个专家标注事件上表现优异。在最贴近实际运维决策的场景中,根因排名质量显著优于规则基线,准确率提升16.1%。尽管训练耗时,但推理仅需数秒GPU时间(远低于典型遥测采集延迟),具备实用部署潜力。

原文摘要 · Abstract (English)

Can a learned model capture how faults propagate through a large-scale network and use this knowledge to causally attribute customer impact to its underlying root cause? Existing root cause analysis techniques often rely on static rules, correlation heuristics, or topology-local reasoning, which struggle to generalize in dynamic environments where faults propagate across complex physical and logical dependencies. We present NetCause, a self-supervised learning-based framework that models network incidents as graph-temporal processes and uses counterfactual simulation to rank candidate root causes. This approach produces an interpretable ranking of root cause hypotheses and integrates naturally with operator-defined mitigation and remediation actions. We train the model on over 1,500 incidents collected over six months from a leading cloud provider's production network and evaluate it on 31 expert-labeled incidents. NetCause consistently improves root cause ranking quality in the regime most relevant to operational decision-making, achieving a 16.1% accuracy improvement over a rule-based heuristic baseline. While training is computationally intensive, inference is lightweight, requiring only seconds of GPU runtime per incident (well below typical telemetry collection latencies).

根因分析图神经网络反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。