arXiv:2412.18298cs.CVcs.AI2024-12被引 16

大模型让视频异常检测更懂语义、会推理、少依赖标注数据。

Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight

  • 用大语言模型和视觉语言模型提升异常的可解释性与语义理解。
  • 实现少样本甚至零样本检测,减少对大量标注数据的依赖。
  • 适合需要开放世界检测与跨场景泛化的实际应用需求。

视频异常检测(VAD)在2024年因引入大语言模型(LLMs)和视觉语言模型(VLMs)取得显著进展,有效应对了动态开放世界场景中的可解释性、时序推理与泛化能力等核心挑战。本文系统综述了基于大模型的前沿方法,聚焦四大方向:(i)通过语义洞察与文本解释增强异常可解释性,使视觉异常更易理解;(ii)捕捉复杂时序关系,实现跨视频帧的动态异常检测与定位;(iii)支持少样本与零样本检测,降低对大规模标注数据的依赖;(iv)结合语义理解与运动特征,处理开放世界及类别无关异常,保障时空一致性。论文强调其重塑VAD范式潜力,并探讨多模态协同优势,提出未来发展方向以充分释放大模型在视频异常检测中的潜力。

原文摘要 · Abstract (English)

Video anomaly detection (VAD) has witnessed significant advancements through the integration of large language models (LLMs) and vision-language models (VLMs), addressing critical challenges such as interpretability, temporal reasoning, and generalization in dynamic, open-world scenarios. This paper presents an in-depth review of cutting-edge LLM-/VLM-based methods in 2024, focusing on four key aspects: (i) enhancing interpretability through semantic insights and textual explanations, making visual anomalies more understandable; (ii) capturing intricate temporal relationships to detect and localize dynamic anomalies across video frames; (iii) enabling few-shot and zero-shot detection to minimize reliance on large, annotated datasets; and (iv) addressing open-world and class-agnostic anomalies by using semantic understanding and motion features for spatiotemporal coherence. We highlight their potential to redefine the landscape of VAD. Additionally, we explore the synergy between visual and textual modalities offered by LLMs and VLMs, highlighting their combined strengths and proposing future directions to fully exploit the potential in enhancing video anomaly detection.

视频异常检测大模型多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。