通过细粒度多模态特征提升视频异常检测精度
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
- 引入细粒度多模态特征,融合视频内容与字幕信息
- 在四个主流数据集上达到当前最佳性能
- 有效减少视觉特征冗余,适合安防场景应用
视频异常检测(VAD)旨在识别连续监控视频中的异常帧。以往方法主要依赖视觉特征的时空相关性来判断视频片段是否存在异常。近期一些工作尝试引入文本等多模态信息以提升效果,但大多以粗粒度方式将文本特征融入视频片段,忽略了视频内部存在的大量冗余信息。为此,本文提出细粒度多模态特征用于视频异常检测(GMFVAD),基于视频片段生成更精细的多模态特征,概括其核心内容,并结合原始视频字幕提取文本特征,增强关键区域的视觉表征。实验表明,所提方法在四个主要数据集上均取得当前最优性能。消融实验证明,性能提升源于冗余信息的有效降低。
原文摘要 · Abstract (English)
Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance the results of video anomaly detection. However, these works merely incorporate text features into video snippets in a coarse manner, overlooking the significant amount of redundant information that may exist within the video snippets. Therefore, we propose to leverage the diversity among multi-modal information to further refine the extracted features, reducing the redundancy in visual features, and we propose Grained Multi-modal Feature for Video Anomaly Detection (GMFVAD). Specifically, we generate more grained multi-modal feature based on the video snippet, which summarizes the main content, and text features based on the captions of original video will be introduced to further enhance the visual features of highlighted portions. Experiments show that the proposed GMFVAD achieves state-of-the-art performance on four mainly datasets. Ablation experiments also validate that the improvement of GMFVAD is due to the reduction of redundant information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。