arXiv:2504.04495cs.CV2025-04被引 15

利用音视频协同提升异常检测鲁棒性,有效降低误报率。

AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection

  • 通过轻量级参数适配实现音视频特征自适应融合,保持预训练模型不变。
  • 在多个数据集上显著提升检测准确率,音视频联合效果优于单一模态。
  • 具备模态缺失时的容错能力,适合复杂真实场景下的智能监控应用。

随着视频异常检测在智能监控中的广泛应用,传统视觉方法在复杂环境下常面临信息不足和高误报问题。本文提出一种新型弱监督框架,利用音视频协同实现鲁棒视频异常检测。基于对比语言-图像预训练(CLIP)在视觉、音频和文本领域的跨模态表征能力,引入两项创新:一是高效的音视频融合机制,通过轻量级参数适配实现自适应跨模态整合,同时保持冻结的CLIP主干;二是动态音视频提示机制,根据音视频特征与文本标签间的语义相关性,增强文本嵌入以提升CLIP在异常检测任务中的泛化能力。此外,为增强推理阶段对模态缺失的鲁棒性,设计了一种基于不确定性的特征蒸馏模块,通过音视频特征多样性建模,动态强化困难特征的表示。实验表明,该框架在多个基准上表现优异,音视频融合显著提升各类场景下的检测精度;即使仅使用视觉输入,经不确定性蒸馏增强后,仍持续优于现有单模态方法。

原文摘要 · Abstract (English)

With the increasing adoption of video anomaly detection in intelligent surveillance domains, conventional visual-based detection approaches often struggle with information insufficiency and high false-positive rates in complex environments. To address these limitations, we present a novel weakly supervised framework that leverages audio-visual collaboration for robust video anomaly detection. Capitalizing on the exceptional cross-modal representation learning capabilities of Contrastive Language-Image Pretraining (CLIP) across visual, audio, and textual domains, our framework introduces two major innovations: an efficient audio-visual fusion that enables adaptive cross-modal integration through lightweight parametric adaptation while maintaining the frozen CLIP backbone, and a novel audio-visual prompt that dynamically enhances text embeddings with key multimodal information based on the semantic correlation between audio-visual features and textual labels, significantly improving CLIP's generalization for the video anomaly detection task. Moreover, to enhance robustness against modality deficiency during inference, we further develop an uncertainty-driven feature distillation module that synthesizes audio-visual representations from visual-only inputs. This module employs uncertainty modeling based on the diversity of audio-visual features to dynamically emphasize challenging features during the distillation process. Our framework demonstrates superior performance across multiple benchmarks, with audio integration significantly boosting anomaly detection accuracy in various scenarios. Notably, with unimodal data enhanced by uncertainty-driven distillation, our approach consistently outperforms current unimodal VAD methods.

异常检测音视频融合多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。