用跨模态融合与注意力机制提升弱监督视频异常检测准确率
Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
- 设计跨模态适配器动态增强音视频相关特征
- 在暴力和裸露数据集上达到最新最佳性能
- 适合关注弱监督视频理解的研究者
弱监督视频异常检测(WS-VAD)近年来成为重要研究方向,仅用视频级标签识别如暴力、裸露等异常事件。但该任务面临模态信息不平衡、正常与异常特征区分困难等挑战。本文提出一种多模态弱监督视频异常检测框架,引入跨模态融合适配器(CFA),动态选择并增强与视觉模态相关的音频-视觉特征;同时提出双曲洛伦兹图注意力机制(HLGAtt),有效捕捉正常与异常表示间的层次关系,提升特征分离精度。大量实验表明,该模型在暴力和裸露检测基准数据集上均达到当前最优效果。
原文摘要 · Abstract (English)
Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this task has substantial challenges, including addressing imbalanced modality information and consistently distinguishing between normal and abnormal features. In this paper, we address these challenges and propose a multi-modal WS-VAD framework to accurately detect anomalies such as violence and nudity. Within the proposed framework, we introduce a new fusion mechanism known as the Cross-modal Fusion Adapter (CFA), which dynamically selects and enhances highly relevant audio-visual features in relation to the visual modality. Additionally, we introduce a Hyperbolic Lorentzian Graph Attention (HLGAtt) to effectively capture the hierarchical relationships between normal and abnormal representations, thereby enhancing feature separation accuracy. Through extensive experiments, we demonstrate that the proposed model achieves state-of-the-art results on benchmark datasets of violence and nudity detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。