用稀疏局部核模型,从高维数据中精准捕捉罕见异常。
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
- 构建自组织稀疏核集成,通过局部建模和竞争机制定位异常。
- 仅用少量核即在数千维空间中识别显著异常点,准确率高。
- 适合科学发现、入侵检测等需解释性与高效性的高维场景。
现代人工智能虽提升了多领域数据表征能力,但其统计特性难以控制,导致异常检测方法常失效。微弱或罕见信号可能隐藏于正常数据的表象规律中,造成检测盲区。本文提出在先验信息极少条件下,异常检测应满足稀疏性、局部性和竞争性三原则。基于此,我们设计SparKer:一种在半监督Neyman-Pearson框架下训练的稀疏高斯核集成,用于局部建模样本与正常参考之间的似然比。理论分析揭示了模型自组织与检测机制。实验证明,该方法在科学发现、开放世界新奇检测、入侵检测及生成模型验证等任务中表现优异。即使仅使用少量核,也能在数千维表示空间中精准定位显著异常,体现其可解释性、效率与可扩展性。
原文摘要 · Abstract (English)
Modern artificial intelligence has revolutionized our ability to extract rich and versatile data representations across scientific disciplines. Yet, the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection (AD) methods to falter. Weak or rare signals can remain hidden within the apparent regularity of normal data, creating a gap in our ability to detect and interpret anomalies. We examine this gap and identify a set of structural desiderata for detection methods operating under minimal prior information: sparsity, to enforce parsimony; locality, to preserve geometric sensitivity; and competition, to promote efficient allocation of model capacity. These principles define a class of self-organizing local kernels that adaptively partition the representation space around regions of statistical imbalance. As an instantiation of these principles, we introduce SparKer, a sparse ensemble of Gaussian kernels trained within a semi-supervised Neyman--Pearson framework to locally model the likelihood ratio between a sample that may contain anomalies and a nominal, anomaly-free reference. We provide theoretical insights into the mechanisms that drive detection and self-organization in the proposed model, and demonstrate the effectiveness of this approach on realistic high-dimensional problems of scientific discovery, open-world novelty detection, intrusion detection, and generative-model validation. Our applications span both the natural- and computer-science domains. We demonstrate that ensembles containing only a handful of kernels can identify statistically significant anomalous locations within representation spaces of thousands of dimensions, underscoring both the interpretability, efficiency and scalability of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。