arXiv:2511.18739cs.AIcs.LG2025-11被引 3

为时序异常检测设计任务导向的评估框架,解决指标选择难问题。

A Problem-Oriented Taxonomy of Evaluation Metrics for Time Series Anomaly Detection

  • 按实际应用场景重新分类20+个评估指标,聚焦任务需求而非数学形式。
  • 发现多数指标能区分异常与噪声,但NAB、Point-Adjust易受随机误报干扰。
  • 适合物联网和工业系统中需精准评估异常检测性能的研究者参考。

时序异常检测广泛应用于物联网与信息物理系统,但其评估因应用目标多样性和指标假设异质而面临挑战。本文提出一种面向问题的框架,基于指标所应对的具体评估难题,而非其数学形式或输出结构,重新诠释现有指标。将二十多个常用指标归纳为六大维度:1)基础准确率评估;2)时效性奖励机制;3)对标注不精确的容忍度;4)反映人工审核成本的惩罚;5)对抗随机或虚高得分的鲁棒性;6)跨数据集基准对比的无参可比性。通过在真实、随机和理想检测场景下进行综合实验,比较各指标得分分布,量化其区分能力——即辨别有效检测与随机噪声的能力。结果表明,尽管多数事件级指标具备强分离性,但若干广泛使用的指标(如NAB、Point-Adjust)对随机得分膨胀表现脆弱。研究揭示指标适用性必须与具体任务强相关,并与物联网应用的运行目标一致。该框架为理解现有指标提供统一分析视角,也为选择或开发更符合上下文、稳健且公平的评估方法提供实践指导。

原文摘要 · Abstract (English)

Time series anomaly detection is widely used in IoT and cyber-physical systems, yet its evaluation remains challenging due to diverse application objectives and heterogeneous metric assumptions. This study introduces a problem-oriented framework that reinterprets existing metrics based on the specific evaluation challenges they are designed to address, rather than their mathematical forms or output structures. We categorize over twenty commonly used metrics into six dimensions: 1) basic accuracy-driven evaluation; 2) timeliness-aware reward mechanisms; 3) tolerance to labeling imprecision; 4) penalties reflecting human-audit cost; 5) robustness against random or inflated scores; and 6) parameter-free comparability for cross-dataset benchmarking. Comprehensive experiments are conducted to examine metric behavior under genuine, random, and oracle detection scenarios. By comparing their resulting score distributions, we quantify each metric's discriminative ability -- its capability to distinguish meaningful detections from random noise. The results show that while most event-level metrics exhibit strong separability, several widely used metrics (e.g., NAB, Point-Adjust) demonstrate limited resistance to random-score inflation. These findings reveal that metric suitability must be inherently task-dependent and aligned with the operational objectives of IoT applications. The proposed framework offers a unified analytical perspective for understanding existing metrics and provides practical guidance for selecting or developing more context-aware, robust, and fair evaluation methodologies for time series anomaly detection.

异常检测评估指标时序数据物联网

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。