arXiv:2409.11223cs.CV2024-09被引 20

融合视觉、运动与音频的多模态注意力机制,提升弱监督异常检测准确率。

Multimodal Attention-Enhanced Feature Fusion-based Weekly Supervised Anomaly Violence Detection

  • 三路特征流:RGB、光流、音频,均通过注意力模块增强时空特征
  • 在三个基准数据集上优于现有方法,显著提升异常检测鲁棒性
  • 适合智能安防领域,尤其对视听异常事件敏感的应用场景

弱监督视频异常检测(WS-VAD)是构建智能监控系统的关键。该系统采用三路特征流:RGB视频、光流与音频信号,每路通过增强注意力模块提取互补的空间与时间特征。第一路基于ViT的CLIP模块进行多阶段特征增强:第一阶段融合Top-k特征与I3D及基于时间上下文聚合(TCA)的丰富时空特征;第二阶段采用不确定性调节双记忆单元(UR-DMU)模型,同时学习正常与异常表征;第三阶段筛选最相关时空特征。第二路利用深度学习与注意力模块从光流中提取增强特征。第三路通过集成VGGish模型的注意力模块捕捉音频线索,以声音模式识别异常。多模态融合整合了运动与音频信息,这些常在纯视觉分析中被忽略。实验在三个基准数据集上验证了该系统的有效性,性能超越现有先进方法。

原文摘要 · Abstract (English)

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream extracts complementary spatial and temporal features using an enhanced attention module to improve detection accuracy and robustness. In the first stream, we employed an attention-based, multi-stage feature enhancement approach to improve spatial and temporal features from the RGB video where the first stage consists of a ViT-based CLIP module, with top-k features concatenated in parallel with I3D and Temporal Contextual Aggregation (TCA) based rich spatiotemporal features. The second stage effectively captures temporal dependencies using the Uncertainty-Regulated Dual Memory Units (UR-DMU) model, which learns representations of normal and abnormal data simultaneously, and the third stage is employed to select the most relevant spatiotemporal features. The second stream extracted enhanced attention-based spatiotemporal features from the flow data modality-based feature by taking advantage of the integration of the deep learning and attention module. The audio stream captures auditory cues using an attention module integrated with the VGGish model, aiming to detect anomalies based on sound patterns. These streams enrich the model by incorporating motion and audio signals often indicative of abnormal events undetectable through visual analysis alone. The concatenation of the multimodal fusion leverages the strengths of each modality, resulting in a comprehensive feature set that significantly improves anomaly detection accuracy and robustness across three datasets. The extensive experiment and high performance with the three benchmark datasets proved the effectiveness of the proposed system over the existing state-of-the-art system.

异常检测多模态融合弱监督视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。