arXiv:2606.29498cs.CV2026-06

提出新方法同时定位异常发生的时间和空间位置。

Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection

论文配图:Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection
图 1 · 摘自论文原文
  • 基于网格块特征,联合建模异常出现的时空位置。
  • 在未标注边界框情况下,实现细粒度空间异常图生成。
  • 在多个数据集上超越现有方法,适合需可解释性的场景。

弱监督视频异常检测(WSVAD)主要关注时间定位,忽略帧内异常的空间范围。但空间定位对可解释性和实际部署至关重要。本文提出一种基于网格块的时空定位框架,联合建模异常发生的时空位置。方法基于网格级块特征,在多实例学习范式下学习区域级异常得分。进一步提出感知邻近关系的Top-k时空选择策略,无需训练时边界框标注即可生成细粒度空间异常图。该方法在多个基准测试中显著提升时空定位精度。此外,我们发布了两个常用数据集测试集的帧级边界框标注,以及代码和预训练模型,为未来空间化弱监督异常检测研究提供新资源。

原文摘要 · Abstract (English)

Weakly supervised video anomaly detection (WSVAD) has predominantly focused on temporal localization, identifying when anomalies occur while largely neglecting their spatial extent within frames. Yet, spatial localization is essential for interpretability and practical deployment in real-world settings. We introduce a patch-based spatiotemporal framework for weakly supervised anomaly localization that jointly models where and when anomalies occur. Our approach operates on grid-level patch features and learns region-level anomaly scores under a multiple instance learning paradigm. We further propose a Proximity-Aware Top-k spatiotemporal selection strategy that enables the model to generate fine-grained spatial anomaly maps without requiring bounding-box supervision during training. Our method surpasses existing state-of-the-art approaches across multiple benchmarks, yielding substantial gains in spatiotemporal localization accuracy. In addition, we release frame-level bounding-box annotations for the test sets of two widely used datasets, along with our code and pretrained models, providing new resources to facilitate future research in spatially grounded WSVAD.

异常检测时空定位弱监督视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。