用评分函数定位异常值根源,兼顾准确与效率。
Score-based Integrated Gradient for Root Cause Explanations of Outliers
- 基于数据似然的评分函数,通过积分梯度追踪异常到正常分布路径
- 在合成图和真实云服务/供应链数据上,优于现有方法的准确率与速度
- 满足多数沙普利值公理,适合高维非线性因果模型的可解释分析
识别异常值的根源是因果推断和异常检测中的基础问题。传统基于启发式或反事实推理的方法在不确定性与高维依赖下表现不佳。本文提出SIREN,一种新颖且可扩展的方法,通过估计数据似然的评分函数来归因异常值根源。归因计算采用积分梯度,沿从异常点到正常数据分布的路径累积评分贡献。该方法满足四个经典沙普利值公理中的三个——零贡献、效率与线性,并满足由底层因果结构导出的不对称公理。与以往工作不同,SIREN直接作用于评分函数,在非线性、高维及异方差因果模型中实现可计算且考虑不确定性的根因归因。在合成随机图以及真实世界云服务和供应链数据集上的大量实验表明,SIREN在归因准确率和计算效率方面均优于当前最先进基线。
原文摘要 · Abstract (English)
Identifying the root causes of outliers is a fundamental problem in causal inference and anomaly detection. Traditional approaches based on heuristics or counterfactual reasoning often struggle under uncertainty and high-dimensional dependencies. We introduce SIREN, a novel and scalable method that attributes the root causes of outliers by estimating the score functions of the data likelihood. Attribution is computed via integrated gradients that accumulate score contributions along paths from the outlier toward the normal data distribution. Our method satisfies three of the four classic Shapley value axioms - dummy, efficiency, and linearity - as well as an asymmetry axiom derived from the underlying causal structure. Unlike prior work, SIREN operates directly on the score function, enabling tractable and uncertainty-aware root cause attribution in nonlinear, high-dimensional, and heteroscedastic causal models. Extensive experiments on synthetic random graphs and real-world cloud service and supply chain datasets show that SIREN outperforms state-of-the-art baselines in both attribution accuracy and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。