用视觉语言模型提升无监督视频异常检测的可解释性与准确率
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
- 结合大模型语义推理能力与选择性提示适配器,聚焦关键语义空间
- 在ShanghaiTech数据集上达到最新最优性能,有效识别长期隐性异常
- 适合需要高可解释性异常检测的工业场景应用
视频异常检测因其在计算机视觉中的关键作用,在学术与工业领域备受关注。然而,异常本身的不可预测性及样本稀缺性给无监督学习带来巨大挑战。为克服无监督方法因缺乏异常先验知识而导致的局限,本文提出VLAVAD(视频-语言模型辅助异常检测)。该方法利用跨模态预训练模型,结合大语言模型的推理能力与选择性提示适配器(SPA)以选取语义空间,并引入序列状态空间模块(S3M)检测语义特征中的时序不一致性。通过将高维视觉特征映射至低维语义空间,显著提升了无监督异常检测的可解释性。所提方法有效应对难以察觉的长期异常检测难题,在具有挑战性的ShanghaiTech数据集上实现当前最优性能。
原文摘要 · Abstract (English)
Video anomaly detection is a subject of great interest across industrial and academic domains due to its crucial role in computer vision applications. However, the inherent unpredictability of anomalies and the scarcity of anomaly samples present significant challenges for unsupervised learning methods. To overcome the limitations of unsupervised learning, which stem from a lack of comprehensive prior knowledge about anomalies, we propose VLAVAD (Video-Language Models Assisted Anomaly Detection). Our method employs a cross-modal pre-trained model that leverages the inferential capabilities of large language models (LLMs) in conjunction with a Selective-Prompt Adapter (SPA) for selecting semantic space. Additionally, we introduce a Sequence State Space Module (S3M) that detects temporal inconsistencies in semantic features. By mapping high-dimensional visual features to low-dimensional semantic ones, our method significantly enhance the interpretability of unsupervised anomaly detection. Our proposed approach effectively tackles the challenge of detecting elusive anomalies that are hard to discern over periods, achieving SOTA on the challenging ShanghaiTech dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。