提出新框架实现弱监督异常检测的精准定位,解决背景干扰与伦理偏见问题。
Localizing to Debias: A Patch-Level Benchmark and Baseline for Weakly Supervised Spatial Anomaly Detection

- 通过动态稀疏化聚焦关键时空区域,抑制背景噪声
- 在三个数据集上实现媲美先进方法的检测性能
- 公开标注与评估协议,支持模型偏见审计
尽管弱监督视频异常检测(WSVAD)日益受到关注,现有方法仍难以弥合粗粒度时间监督与细粒度空间推理之间的差距。一个关键障碍是时间检测器倾向于依赖背景和场景级线索,而非真正具有判别性的异常证据。这种背景偏差引发伦理问题:模型可能将异常错误关联于社会或环境上下文,而非真实的犯罪相关线索。由于缺乏空间定位,此类偏差难以被发现和审计。为此,我们提出SST-WSVADL,一种稀疏时空框架,将时间异常检测与细粒度空间定位相衔接。该框架不盲目处理所有空间区域,而是通过动态稀疏化逐步聚焦最具异常相关性的时空区域,自然抑制主导性背景内容并保留判别性证据。时间与空间分支通过运动感知正则化端到端耦合,引导稀疏化向动态信息丰富区域靠拢,无需外部检测器或视觉语言提示。我们公开了帧级空间标注及适用于UCF-Crime、XD-Violence和MSAD三个公开数据集的方法无关评估协议,使社区能够审计WSVAD预测中的空间偏差,推动更伦理、可问责的异常检测发展。实验表明,SST-WSVADL在多个基准上表现优异,同时支持异常定位与块级别偏差审计,为面向可解释性的评估提供可复现基础。
原文摘要 · Abstract (English)
Despite growing interest in weakly supervised video anomaly detection (WSVAD), current methods struggle to bridge the gap between coarse temporal supervision and fine-grained spatial reasoning. A key obstacle is the tendency of temporal detectors to latch onto background and scene-level cues rather than truly discriminative anomaly evidence. This background bias raises ethical concerns: models may inadvertently associate anomalies with societal or environmental context rather than authentic crime-related cues. Without spatial grounding, such biases remain hidden and unauditable. To address this, we propose SST-WSVADL, a sparse spatio-temporal framework that bridges temporal anomaly detection with fine-grained spatial localization. Rather than processing all spatial regions indiscriminately, SST-WSVADL progressively focuses on the most anomaly-relevant spatio-temporal regions through dynamic sparsification, naturally suppressing background dominant content while preserving discriminative evidence. The temporal and spatial branches are coupled end-to-end via motion-aware regularization that guides sparsification toward dynamically informative regions, without relying on external detectors or vision-language prompts. We publicly release frame-level spatial annotations and a method-agnostic evaluation protocol for three public datasets: UCF-Crime, XD-Violence, and MSAD. These resources enable the community to audit spatial biases in WSVAD predictions, supporting progress toward more ethical and accountable anomaly detection. Experiments demonstrate that SST-WSVADL is competitive with prior methods across benchmarks while enabling localization and patch-level auditability of scene bias, providing a reproducible foundation for interpretability-oriented evaluation of WSVAD models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。