提升弱监督视频异常检测的事件定位完整性
Learning Event Completeness for Weakly Supervised Video Anomaly Detection
- 设计双结构模型,融合视觉与语言的类别相关和无关语义
- 通过异常感知高斯混合模型精确捕捉异常事件边界
- 用记忆库原型学习增强文本描述,适合研究异常检测的学者
弱监督视频异常检测(WS-VAD)旨在仅使用视频级标注的情况下,定位未剪裁视频中包含异常事件的时间段。然而,由于缺乏密集帧级标注,现有方法常出现定位不完整的问题。为此,本文提出LEC-VAD,一种基于事件完整性学习的弱监督视频异常检测新框架。该框架采用双结构设计,以编码视觉与语言间的类别感知和类别无关语义。在模型内部,利用异常感知高斯混合模型构建语义规律,实现对异常事件边界的精准学习,从而生成更完整的事件实例。此外,设计了一种基于记忆库的原型学习机制,丰富与异常事件类别相关的简短文本描述,显著提升文本表达能力,对推动WS-VAD发展具有重要意义。在两个基准数据集XD-Violence和UCF-Crime上,LEC-VAD均显著优于当前最先进方法。
原文摘要 · Abstract (English)
Weakly supervised video anomaly detection (WS-VAD) is tasked with pinpointing temporal intervals containing anomalous events within untrimmed videos, utilizing only video-level annotations. However, a significant challenge arises due to the absence of dense frame-level annotations, often leading to incomplete localization in existing WS-VAD methods. To address this issue, we present a novel LEC-VAD, Learning Event Completeness for Weakly Supervised Video Anomaly Detection, which features a dual structure designed to encode both category-aware and category-agnostic semantics between vision and language. Within LEC-VAD, we devise semantic regularities that leverage an anomaly-aware Gaussian mixture to learn precise event boundaries, thereby yielding more complete event instances. Besides, we develop a novel memory bank-based prototype learning mechanism to enrich concise text descriptions associated with anomaly-event categories. This innovation bolsters the text's expressiveness, which is crucial for advancing WS-VAD. Our LEC-VAD demonstrates remarkable advancements over the current state-of-the-art methods on two benchmark datasets XD-Violence and UCF-Crime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。