首个统一评估弱监督异常检测的基准,揭示不同标签缺陷间的内在关联。
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark

- 构建统一评测框架,系统对比36种算法在四种模态下的表现。
- 发现专用模型仅在极低标签量下占优,其余场景被通用模型超越。
- 揭示标签噪声类型对模型影响不对称,无标签数据增益有限。
弱监督异常检测(WSAD)发展出不完整、不精确和不准确三种主要方向,但彼此孤立,缺乏统一评估框架以判断其应对独特挑战还是共享底层机制。本文提出首个统一评测基准WSADBench,涵盖4种模态,标准化评估36种算法,通过系统性地改变标签数量、粒度与质量,开展超过70万次实验。结果揭示四大关键发现:(i) 这些弱监督场景间存在强内在关联,挑战现有研究方向的独立性;(ii) 专用WSAD算法仅在极端标签稀缺条件下表现优异,随着标签增多或在分布外(OOD)场景中,迅速被表格基础模型和通用分类方法超越;(iii) 无标签数据在不同设置下效用不一,相比标签精炼带来的提升微弱;(iv) 模型对不同类型的标签噪声呈现非对称敏感性。相关代码与数据集已开源,网址:https://github.com/SUFE-AILAB/WSADBench。
原文摘要 · Abstract (English)
Weakly supervised anomaly detection (WSAD) has developed in three primary directions: incomplete, inexact, and inaccurate supervision. However, these directions remain isolated, lacking a unified framework to assess whether they address unique challenges or share fundamental mechanisms. This paper introduces WSADBench, the first benchmark that unifies evaluation across distinct weakly supervised scenarios, benchmarking diverse approaches from specialized WSAD methods to advanced tabular foundation models. WSADBench establishes standardized protocols to evaluate 36 algorithms across 4 modalities by systematically varying label quantity, granularity, and quality, revealing the performance boundaries of various methods. Based on over 700K experiments, WSADBench reveals four critical insights: (i) Strong intrinsic correlations exist between these weak supervision scenarios, challenging the isolation of current research directions. (ii) Specialized WSAD algorithms excel only in extreme label-scarcity regimes but are quickly dominated by tabular foundation models and general classification methods as supervision increases or in OOD scenarios. (iii) Unlabeled data shows inconsistent utility across settings, with marginal gains compared to label refinement. (iv) Models exhibit asymmetric sensitivity to different types of label noise. We release WSADBench as an open-source benchmark with code and datasets to facilitate future WSAD research: https://github.com/SUFE-AILAB/WSADBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。