为聚类异常检测提供统计推断框架,确保误报率可控。
Statistical Inference for Clustering-based Anomaly Detection
- 基于选择性推断理论,分析聚类异常检测的选中机制。
- 在显著性水平α=0.05下严格控制误报率。
- 提升真实异常检出率,适合对可靠性要求高的场景。
无监督异常检测是机器学习与统计学中的基础问题。聚类驱动的异常检测方法虽常用,但缺乏对检测结果可靠性的保障。本文提出SI-CLAD(Statistical Inference for CLustering-based Anomaly Detection),一种用于检验聚类异常检测结果的新统计框架。其核心优势在于可严格控制虚假异常识别的概率,使其低于预设显著性水平α(如α=0.05)。通过分析聚类异常检测中固有的选择机制,并结合选择性推断(Selective Inference, SI)框架,我们证明了误报控制的可行性。此外,引入策略提升真实异常检出率,增强整体性能。在合成数据与真实数据集上的大量实验验证了理论结果,展示了该方法的优越性。
原文摘要 · Abstract (English)
Unsupervised anomaly detection (AD) is a fundamental problem in machine learning and statistics. A popular approach to unsupervised AD is clustering-based detection. However, this method lacks the ability to guarantee the reliability of the detected anomalies. In this paper, we propose SI-CLAD (Statistical Inference for CLustering-based Anomaly Detection), a novel statistical framework for testing the clustering-based AD results. The key strength of SI-CLAD lies in its ability to rigorously control the probability of falsely identifying anomalies, maintaining it below a pre-specified significance level $α$ (e.g., $α= 0.05$). By analyzing the selection mechanism inherent in clustering-based AD and leveraging the Selective Inference (SI) framework, we prove that false detection control is attainable. Moreover, we introduce a strategy to boost the true detection rate, enhancing the overall performance of SI-CLAD. Extensive experiments on synthetic and real-world datasets provide strong empirical support for our theoretical findings, showcasing the superior performance of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。