发现表格数据中异常样本罕见被赋予高似然,与图像领域相反。
Why the Counterintuitive Phenomenon of Likelihood Rarely Appears in Tabular Anomaly Detection with Deep Generative Models?
- 提出统一评估框架,定义并验证该反直觉现象在表格数据中极少出现。
- 在47个表格数据集上测试,13种基线模型均显示异常样本似然普遍较低。
- 适用于需要可靠异常检测的金融、医疗等表格数据场景。
基于可解析计算似然的深度生成模型(如归一化流)可通过似然评分实现异常检测。我们发现,与图像领域中异常样本常获更高似然不同,该反直觉现象在表格数据中极为罕见。本文提出一种无领域依赖的统一评估范式,解决该现象缺乏明确定义的问题。在ADBench的47个表格数据集和10个CV/NLP嵌入数据集上,对比13种基线模型的广泛实验表明,按定义衡量,该现象在一般表格数据中始终稀少。从理论与实证双重视角分析,结果表明数据维度与特征相关性差异是关键影响因素。研究证实,仅依赖似然的归一化流方法在表格异常检测中具有实际可行性与可靠性。
原文摘要 · Abstract (English)
Deep generative models with tractable and analytically computable likelihoods, exemplified by normalizing flows, offer an effective basis for anomaly detection through likelihood-based scoring. We demonstrate that, unlike in the image domain where deep generative models frequently assign higher likelihoods to anomalous data, such counterintuitive behavior occurs far less often in tabular settings. We first introduce a domain-agnostic formulation that enables consistent detection and evaluation of the counterintuitive phenomenon, addressing the absence of precise definition. Through extensive experiments on 47 tabular datasets and 10 CV/NLP embedding datasets in ADBench, benchmarked against 13 baseline models, we demonstrate that the phenomenon, as defined, is consistently rare in general tabular data. We further investigate this phenomenon from both theoretical and empirical perspectives, focusing on the roles of data dimensionality and difference in feature correlation. Our results suggest that likelihood-only detection with normalizing flows offers a practical and reliable approach for anomaly detection in tabular domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。