arXiv:2502.08593cs.LG2025-02中稿 · UAI 2025被引 3

用信息论统一异常传播规律,可定位异常根源。

Toward Universal Laws of Outlier Propagation

  • 基于算法因果网络,将联合样本的随机性缺陷分解为各机制贡献之和
  • 弱异常无法引发强异常,揭示异常传播的因果限制
  • 适用于异常归因与检测,尤其适合理解已有评分方法

当多种异常特征导致不同样本被标记为异常时,算法信息论(AIT)提供了一种统一的框架,以样本的随机性缺陷来刻画异常。在因果贝叶斯网络的算法马尔可夫条件下,我们证明联合样本的随机性缺陷可分解为各因果机制上随机性缺陷之和。因此,异常观测可归因于其根本原因——即行为异常的机制。作为莱文随机性守恒定律的扩展,我们进一步证明弱异常无法引发强异常。这些信息论法则有助于澄清异常检测与归因的理解,并与先前文献中的特定异常评分方法相呼应。

原文摘要 · Abstract (English)

When a variety of anomalous features motivate flagging different samples as outliers, Algorithmic Information Theory (AIT) offers a principled way to unify them in terms of a sample's randomness deficiency. Subject to the algorithmic Markov condition on a causal Bayesian network, we show that the randomness deficiency of a joint sample decomposes into a sum of randomness deficiencies at each causal mechanism. Consequently, anomalous observations can be attributed to their root causes, i.e., the mechanisms that behaved anomalously. As an extension of Levin's law of randomness conservation, we show that weak outliers cannot cause strong ones. We show how these information theoretic laws clarify our understanding of outlier detection and attribution, in the context of more specialized outlier scores from prior literature.

异常检测因果推理信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。