为深度近邻异常检测提供可解释的显著性评估,控制误报率。
Quantifying Statistical Significance of Deep Nearest Neighbor Anomaly Detection via Selective Inference
- 基于选择性推断,对深度嵌入空间中的近邻距离进行统计检验。
- 输出异常检测的p值,可在0.05等水平上严格控制误报率。
- 适用于工业质检等需高可靠性异常检测的场景。
在真实应用中,异常检测常缺乏异常数据,需依赖仅含正常样本的半监督方法。其中,深度k近邻(deep kNN)异常检测因其可解释性和灵活性脱颖而出,通过深度隐空间中的距离评分实现检测。尽管性能优异,deep kNN缺乏不确定性量化机制,这在工业检测等关键场景中尤为不足。为此,我们提出一种统计框架,以p值形式量化检测结果的显著性,从而在用户指定的显著性水平(如0.05)下控制假阳性率。核心挑战在于处理由数据驱动选择带来的选择偏差,我们采用选择性推断这一严谨方法,在条件选择基础上进行统计推断。我们在多种数据集上评估该方法,证明其在工业应用场景中具备可靠的异常检测能力。
原文摘要 · Abstract (English)
In real-world applications, anomaly detection (AD) often operates without access to anomalous data, necessitating semi-supervised methods that rely solely on normal data. Among these methods, deep k-nearest neighbor (deep kNN) AD stands out for its interpretability and flexibility, leveraging distance-based scoring in deep latent spaces.Despite its strong performance, deep kNN lacks a mechanism to quantify uncertainty-an essential feature for critical applications such as industrial inspection. To address this limitation, we propose a statistical framework that quantifies the significance of detected anomalies in the form of p-values, thereby enabling control over false positive rates at a user-specified significance level (e.g.,0.05). A central challenge lies in managing selection bias, which we tackle using Selective Inference-a principled method for conducting inference conditioned on data-driven selections. We evaluate our method on diverse datasets and demonstrate that it provides reliable AD well-suited for industrial use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。