arXiv:2507.15584cs.LG2025-07被引 4

现有异常检测评测标准有缺陷,需按应用场景重构评估体系。

We Need to Rethink Benchmarking in Anomaly Detection

  • 提出按应用特征分场景评测,替代通用基准测试
  • 发现仅检测极端值的简单算法能媲美深度学习模型
  • 适合关注实际应用选型的研究者和从业者

尽管不断有新异常检测算法提出且评测工作持续开展,但性能提升却趋于停滞,现有基线与新方法差异微小。本文认为这种停滞源于评估方式的局限:当前基准测试中,仅检查单个特征极端值的简单算法,竟可与先进深度学习方法表现相当,却无法处理环形正常区域中的异常等基础情形。此外,现有评测未能涵盖异常检测应用的多样性,使实践者难以可靠选择适用算法。因此,我们主张重新思考异常检测的评测范式——应基于共性特征将应用划分为不同场景,通过场景内定制预处理、指标与模型选择,明确哪些进展可跨场景迁移,为实际应用提供可信指导。

原文摘要 · Abstract (English)

Despite the continuous proposal of new anomaly detection algorithms and extensive benchmarking efforts, progress seems to stagnate, with only minor performance differences between established baselines and new algorithms. In this position paper, we argue that this stagnation is due to limitations in how we evaluate anomaly detection algorithms. In current benchmarks, a trivial algorithm that only checks for extreme values in individual features performs competitively with state-of-the-art deep learning methods, despite failing on simple cases such as anomalies within an annulus of normal points. Moreover, existing benchmarks do not adequately reflect the diversity of anomaly detection applications, making it difficult for practitioners to reliably select algorithms for their applications. Consequently, we need to rethink benchmarking in anomaly detection. In our opinion, anomaly detection should be studied using scenarios that group applications sharing relevant characteristics, defined through a common taxonomy. Benchmarking within scenarios enables scenario-specific choices for preprocessing, metrics, and model selection, clarifying which advances transfer across similar applications and providing practitioners with reliable guidance for their specific contexts.

异常检测评测基准场景划分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。