arXiv:2502.18002cs.LGcs.AI2025-02

用测度论中的辐射-尼科迪姆导数改进异常检测损失函数,提升性能。

A Radon-Nikodým Perspective on Anomaly Detection: Theory and Implications

  • 引入辐射-尼科迪姆导数修正损失函数,统一监督与无监督场景
  • 在96个数据集上,多变量数据中68%的F1分数超越现有方法
  • 适用于医疗、金融、网络安全等领域的时序与多变量异常检测

什么原则能指导有效的异常检测损失函数设计?答案来自测度论中的辐射-尼科迪姆定理。本文核心观点:将原始损失函数乘以辐射-尼科迪姆导数,可全面提升性能,称为RN-Loss。我们基于PAC可学习性框架证明该方法的有效性。根据上下文,辐射-尼科迪姆导数呈现不同形式:在监督异常检测中为加权损失;在无监督场景(带分布假设)下则对应流行的基于聚类的局部离群因子。我们在96个数据集上评估该算法,涵盖医疗、网络安全、金融等领域的单变量与多变量数据。结果表明,在多变量数据中,RN-Loss在68%的数据集上优于现有最先进方法(以F1分数衡量),并在72%的时间序列(单变量)数据集上达到峰值F1分数。

原文摘要 · Abstract (English)

Which principle underpins the design of an effective anomaly detection loss function? The answer lies in the concept of Radon-Nikodým theorem, a fundamental concept in measure theory. The key insight from this article is: Multiplying the vanilla loss function with the Radon-Nikodým derivative improves the performance across the board. We refer to this as RN-Loss. We prove this using the setting of PAC (Probably Approximately Correct) learnability. Depending on the context a Radon-Nikodým derivative takes different forms. In the simplest case of supervised anomaly detection, Radon-Nikodým derivative takes the form of a simple weighted loss. In the case of unsupervised anomaly detection (with distributional assumptions), Radon-Nikodým derivative takes the form of the popular cluster based local outlier factor. We evaluate our algorithm on 96 datasets, including univariate and multivariate data from diverse domains, including healthcare, cybersecurity, and finance. We show that RN-Derivative algorithms outperform state-of-the-art methods on 68% of Multivariate datasets (based on F1 scores) and also achieves peak F1-scores on 72% of time series (Univariate) datasets.

异常检测测度论损失函数机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。