用五种新模态增强视频异常检测,提升真实场景下识别精度。
Just Dance with $π$! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
- 融合姿态、深度、全景掩码等五种模态,强化RGB特征
- 在三个真实场景数据集上达到最新最好效果
- 训练时用多模态,推理时仅需原始视频,效率高
弱监督视频异常检测(VAD)传统上仅依赖RGB时空特征,难以区分如盗窃与外观相似事件,限制了其在真实场景中的可靠性。为此,本文提出多模态诱导框架PI-VAD,通过引入五种额外模态——姿态(Pose)、三维场景与实体表示(Depth)、周围物体(Panoptic masks)、全局运动(光流)及语言提示(VLM)——增强RGB表征。每个模态构成多边形的一条边,协同提升特征显著性。PI-VAD包含两个可插拔模块:伪模态生成模块与跨模态诱导模块,分别生成特定模态原型表示,并将多模态信息注入RGB特征中。这两模块通过异常感知的辅助任务实现,需五个模态骨干网络,但仅用于训练。值得注意的是,该方法在三个典型真实场景VAD数据集上达到当前最优性能,且推理阶段无需五模态骨干网络,无额外计算开销。
原文摘要 · Abstract (English)
Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such as shoplifting from visually similar events. Therefore, towards robust complex real-world VAD, it is essential to augment RGB spatio-temporal features by additional modalities. Motivated by this, we introduce the Poly-modal Induced framework for VAD: "PI-VAD", a novel approach that augments RGB representations by five additional modalities. Specifically, the modalities include sensitivity to fine-grained motion (Pose), three dimensional scene and entity representation (Depth), surrounding objects (Panoptic masks), global motion (optical flow), as well as language cues (VLM). Each modality represents an axis of a polygon, streamlined to add salient cues to RGB. PI-VAD includes two plug-in modules, namely Pseudo-modality Generation module and Cross Modal Induction module, which generate modality-specific prototypical representation and, thereby, induce multi-modal information into RGB cues. These modules operate by performing anomaly-aware auxiliary tasks and necessitate five modality backbones -- only during training. Notably, PI-VAD achieves state-of-the-art accuracy on three prominent VAD datasets encompassing real-world scenarios, without requiring the computational overhead of five modality backbones at inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。