arXiv:2411.14515cs.LGcs.AI2024-11被引 3

提出多层级异常检测新范式,让分数更真实反映异常严重程度

Are Anomaly Scores Telling the Whole Story? A Benchmark for Multilevel Anomaly Detection

  • 设计多层级异常检测(MAD)新任务,让评分体现真实严重性
  • 构建MAD-Bench基准,评估模型在严重性判断上的表现
  • 揭示现有模型在严重性对齐上的不足,适合实际部署研究者参考

异常检测(AD)通过学习正常数据模式来识别异常。在真实场景中,异常严重程度各异,从轻微问题到需紧急处理的严重异常不等。但现有模型多为二值判断,其异常分数仅反映与正常数据的偏离程度,难以准确体现实际严重性。本文提出多层级异常检测(MAD)新设定,使异常分数直接对应真实严重程度,并强调其在多个领域的应用价值。我们构建了MAD-Bench基准,不仅评估模型的异常检出能力,更关注其评分是否与严重性对齐,包含多种基线和涉及严重性的现实应用场景。我们在该基准上进行综合性能分析,评估模型分配与严重性一致分数的能力,探究二值检测与多层级检测之间的表现关联,并研究模型鲁棒性。分析结果为提升模型在真实场景中的严重性对齐能力提供了关键洞见。代码框架与数据集将公开共享。

原文摘要 · Abstract (English)

Anomaly detection (AD) is a machine learning task that identifies anomalies by learning patterns from normal training data. In many real-world scenarios, anomalies vary in severity, from minor anomalies with little risk to severe abnormalities requiring immediate attention. However, existing models primarily operate in a binary setting, and the anomaly scores they produce are usually based on the deviation of data points from normal data, which may not accurately reflect practical severity. In this paper, we address this gap by making three key contributions. First, we propose a novel setting, Multilevel AD (MAD), in which the anomaly score represents the severity of anomalies in real-world applications, and we highlight its diverse applications across various domains. Second, we introduce a novel benchmark, MAD-Bench, that evaluates models not only on their ability to detect anomalies, but also on how effectively their anomaly scores reflect severity. This benchmark incorporates multiple types of baselines and real-world applications involving severity. Finally, we conduct a comprehensive performance analysis on MAD-Bench. We evaluate models on their ability to assign severity-aligned scores, investigate the correspondence between their performance on binary and multilevel detection, and study their robustness. This analysis offers key insights into improving AD models for practical severity alignment. The code framework and datasets used for the benchmark will be made publicly available.

异常检测严重性评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。