arXiv:2608.11985cs.CV2026-08中稿 · ECCV被引 3

揭示弱监督视频异常检测中帧级AUC的评估缺陷,指出其无法可靠衡量定位能力。

Auditing Frame-Level AUC in Weakly Supervised Video Anomaly Detection: Granularity, Resolution, and Scene Bias

论文配图:Auditing Frame-Level AUC in Weakly Supervised Video Anomaly Detection: Granularity, Resolution, and Scene Bias
图 1 · 摘自论文原文
  • 对比全局、类别级、视频内三种粒度下的帧排序,发现标准AUC依赖视频间差异
  • 在UCF-Crime数据集上,标准AUC无法分辨顶尖模型间的微小差距,而视频内AUC可
  • 模型对视频分辨率、色彩编码等场景因素敏感,说明偏差来自数据而非架构

帧级受试者工作特征曲线下面积(AUC)是弱监督视频异常检测(WSVAD)的主要评估指标。其标准形式衡量异常帧是否优于任意测试视频中的正常帧,我们称之为池化AUC,因其将不同视频的帧对混合比较。池化AUC同时反映事件定位能力与视频间差异。我们在UCF-Crime数据集上对近期多种骨干网络的先进模型进行审计,固定模型帧得分,在三种配对粒度下重读:全局、按异常类别、每段视频内部;并重复此过程于零样本内部表征计算得分。通过成对视频自助法评估排序可靠性。结果发现:第一,池化AUC不能可靠预测视频内定位性能,相似池化得分模型在更严格粒度下表现差异大且排名反转;第二,在基准测试集规模下,池化AUC分辨率不足,无法支持领域报告的顶尖差距,同一骨干族内无法区分,而视频内AUC可区分相同预测;第三,仅在正常视频上,所有模型均按拍摄属性(如分辨率、色彩编码)区分视频,表明场景敏感性普遍存在于该任务,非特定于某架构。我们公开发布一个可从现有预测和场景因子标注计算的粒度感知评估协议。

原文摘要 · Abstract (English)

Frame-level area under the ROC curve (AUC) is the dominant evaluation metric for weakly supervised video anomaly detection (WSVAD). Its standard form measures whether an anomalous frame outranks a normal frame drawn from anywhere in the test set. We refer to this comparison as pooled AUC, since it aggregates frame pairs across test videos regardless of source. Pooled AUC therefore credits both event localization and differences between video sources. We audit this protocol on UCF-Crime across recent state-of-the-art models spanning different backbone families. Holding each model's frame scores fixed, we read them under three pairing granularities: global, per anomaly category, and within each video, then repeat the same three-granularity readout on zero-shot scores computed from the models' internal representations. We assess ranking reliability with a paired video bootstrap. Three findings follow. First, pooled AUC does not reliably predict within-video anomaly localization: models with similar pooled scores exhibit large localization differences and rank reversals under stricter granularities. Second, at the benchmark's test-split size, pooled AUC lacks the resolution to support state-of-the-art margins reported in the field. Within each backbone family, it resolves no comparison at those margins, while within-video AUC resolves several over identical predictions. Learned representations further reveal that within-video anomaly structure and detector localization are decoupled. Third, on normal footage alone, every model we examine separates videos by recording properties, such as resolution and color encoding, indicating that scene sensitivity is shared across the setting rather than specific to any architecture. We publicly release a granularity-aware protocol computable from existing predictions and scene-factor annotations for UCF-Crime.

视频异常检测评估漏洞粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。