arXiv:2603.06131cs.LG2026-03被引 1

提出新评估指标DQE,更准确衡量时间序列异常检测效果

DQE: A Semantic-Aware Evaluation Metric for Time Series Anomaly Detection

  • 基于检测语义划分异常区域为三类子区域
  • 在合成与真实数据上优于10种常用评估指标
  • 适合评估异常检测模型的稳定性与可解释性

近年来时间序列异常检测取得了显著进展,但评估方法仍缺乏足够关注。现有指标存在四大局限:(1) 偏向点级覆盖率,(2) 对接近异常的检测不敏感或不一致,(3) 对误报惩罚不足,(4) 受阈值或阈值区间选择影响导致结果不一致。这些缺陷可能导致不可靠或反直觉的评估结果,阻碍客观进步。本文从检测语义出发,重新审视评估问题,提出DQE指标。通过引入基于语义的区域划分策略,将每个异常的局部时间区间分解为三个功能不同的子区域,进而设计针对各子区域的细粒度评分机制,实现更可靠、可解释的综合评估。通过系统分析现有指标,发现阈值区间选择带来的评估偏差,并采用全阈值谱聚合策略消除不一致性。在合成数据与真实数据上的大量实验表明,DQE能提供稳定、区分度高且可解释的评估结果,相比十种常用指标具有更强鲁棒性。

原文摘要 · Abstract (English)

Time series anomaly detection has achieved remarkable progress in recent years. However, evaluation practices have received comparatively less attention, despite their critical importance. Existing metrics exhibit several limitations: (1) bias toward point-level coverage, (2) insensitivity or inconsistency in near-miss detections, (3) inadequate penalization of false alarms, and (4) inconsistency caused by threshold or threshold-interval selection. These limitations can produce unreliable or counterintuitive results, hindering objective progress. In this work, we revisit the evaluation of time series anomaly detection from the perspective of detection semantics and propose a novel metric for more comprehensive assessment. We first introduce a partitioning strategy grounded in detection semantics, which decomposes the local temporal region of each anomaly into three functionally distinct subregions. Using this partitioning, we evaluate overall detection behavior across events and design finer-grained scoring mechanisms for each subregion, enabling more reliable and interpretable assessment. Through a systematic study of existing metrics, we identify an evaluation bias associated with threshold-interval selection and adopt an approach that aggregates detection qualities across the full threshold spectrum, thereby eliminating evaluation inconsistency. Extensive experiments on synthetic and real-world data demonstrate that our metric provides stable, discriminative, and interpretable evaluation, while achieving robust assessment compared with ten widely used metrics.

异常检测评估指标时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。