无需已知因果图,通过局部机制变化定位系统故障根源。
StableRCA: Robust Graph-Agnostic Mechanism-Level Root Cause Analysis

- 基于局部马尔可夫边界和分布漂移检测,不依赖全局因果图。
- 在多种真实数据集上表现稳定,支持多目标干预且样本量增大时准确率指数提升。
- 适合复杂系统故障诊断,尤其适用于因果结构未知的场景。
根因分析(RCA)旨在识别制造、云计算和医疗等复杂领域中异常系统行为的根源变量。现有方法存在瓶颈:基于图的因果方法需已知或准确估计因果图,而无图统计方法要么仅定位边缘异常,要么依赖对图结构或函数形式的严苛假设。本文提出StableRCA,一种局部机制级的根因分析框架,通过估计局部马尔可夫边界并检测其内部条件分布漂移来避免全局图发现。基于独立因果机制原理,证明在马尔可夫边界恢复准确且机制漂移非退化条件下,干预目标的识别概率随样本量呈指数收敛。在合成基准与五个真实数据集上的实验表明,StableRCA对图误设具有鲁棒性,支持多干预目标,可扩展至大规模系统,并在不同应用领域均表现可靠。代码已公开:https://anonymous.4open.science/r/StableRCA-E362。
原文摘要 · Abstract (English)
Root-Cause Analysis (RCA) seeks to identify the variables responsible for abnormal system behavior in complex domains such as manufacturing, cloud computing, and healthcare. Existing approaches face a critical bottleneck: graph-based causal methods can identify intervention targets but typically require a known or accurately estimated causal graph, while graph-free statistical methods either localize marginal anomalies rather than structural causes, or rely on restrictive assumptions about graph structure or functional form. We propose StableRCA, a local mechanism-level RCA framework that avoids global graph discovery by estimating local Markov boundaries and detecting conditional distribution shifts within them. Leveraging the Independent Causal Mechanism principle, we show that intervention targets can be identified with probability converging exponentially in sample size under faithful Markov boundary recovery and non-degenerate mechanism shifts. Experiments on synthetic benchmarks and five real-world datasets demonstrate that StableRCA is robust to graph misspecification, effective under multiple intervention targets, scalable to large systems, and reliable across diverse application domains. Code is available at: https://anonymous.4open.science/r/StableRCA-E362
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。