用文本指导提升弱监督视频异常检测效果
Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
- 通过上下文学习生成高质量异常文本用于模型微调
- 多尺度瓶颈注意力融合有效缓解模态冗余与不平衡
- 适合关注视频异常检测与多模态融合的研究者
弱监督多模态视频异常检测受到广泛关注,但文本模态的潜力尚未充分挖掘。文本提供明确语义信息,有助于增强异常表征并减少误报。然而,通用语言模型难以捕捉异常特异性细节,且缺乏相关描述。此外,多模态融合常面临冗余与失衡问题。为此,我们提出一种新型文本引导框架:首先,设计基于上下文学习的多阶段文本增强机制,生成高质量异常文本样本以微调文本特征提取器;其次,构建多尺度瓶颈Transformer融合模块,利用压缩后的瓶颈令牌逐步融合跨模态信息,缓解冗余与不平衡。在UCF-Crime和XD-Violence数据集上的实验表明,该方法达到领先性能。
原文摘要 · Abstract (English)
Weakly supervised multimodal video anomaly detection has gained significant attention, yet the potential of the text modality remains under-explored. Text provides explicit semantic information that can enhance anomaly characterization and reduce false alarms. However, extracting effective text features is challenging due to the inability of general-purpose language models to capture anomaly-specific nuances and the scarcity of relevant descriptions. Furthermore, multimodal fusion often suffers from redundancy and imbalance. To address these issues, we propose a novel text-guided framework. First, we introduce an in-context learning-based multi-stage text augmentation mechanism to generate high-quality anomaly text samples for fine-tuning the text feature extractor. Second, we design a multi-scale bottleneck Transformer fusion module that uses compressed bottleneck tokens to progressively integrate information across modalities, mitigating redundancy and imbalance. Experiments on UCF-Crime and XD-Violence demonstrate state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。