首篇系统综述多模态大模型时代的视频异常检测方法演化。
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- 构建涵盖深度网络与大模型的统一框架,梳理技术演进路径。
- 分析多模态大模型对标注、输入、任务目标的变革性影响。
- 适合关注智能监控与大模型融合应用的研究者参考。
视频异常检测(VAD)旨在识别并定位视频中的异常行为或事件,是智能监控与公共安全领域的核心技术。随着深度学习的发展,深度模型架构的持续演进推动了VAD方法的创新,显著提升了特征表示能力与场景适应性,增强了算法泛化性并拓展了应用边界。更重要的是,多模态大语言模型(MLLMs)和大语言模型(LLMs)的快速发展为VAD带来了新机遇与挑战。在MLLMs和LLMs支持下,VAD在数据标注、输入模态、模型架构与任务目标等方面发生显著变革。近年来论文数量激增与任务演变催生了系统性综述的迫切需求。本文首次全面分析基于MLLMs和LLMs的VAD方法,深入探讨大模型时代VAD领域的变迁及其成因。提出一个涵盖深度神经网络(DNN)与大语言模型(LLM)的统一框架,对大模型赋能的新范式进行系统分析,构建分类体系,并对比其优劣。在此基础上,聚焦当前基于MLLMs/LLMs的VAD方法。最后,结合技术演进轨迹与现存瓶颈,提炼关键挑战并展望未来研究方向,为该领域提供指导。
原文摘要 · Abstract (English)
Video anomaly detection (VAD) aims to identify and ground anomalous behaviors or events in videos, serving as a core technology in the fields of intelligent surveillance and public safety. With the advancement of deep learning, the continuous evolution of deep model architectures has driven innovation in VAD methodologies, significantly enhancing feature representation and scene adaptability, thereby improving algorithm generalization and expanding application boundaries. More importantly, the rapid development of multi-modal large language (MLLMs) and large language models (LLMs) has introduced new opportunities and challenges to the VAD field. Under the support of MLLMs and LLMs, VAD has undergone significant transformations in terms of data annotation, input modalities, model architectures, and task objectives. The surge in publications and the evolution of tasks have created an urgent need for systematic reviews of recent advancements. This paper presents the first comprehensive survey analyzing VAD methods based on MLLMs and LLMs, providing an in-depth discussion of the changes occurring in the VAD field in the era of large models and their underlying causes. Additionally, this paper proposes a unified framework that encompasses both deep neural network (DNN)-based and LLM-based VAD methods, offering a thorough analysis of the new VAD paradigms empowered by LLMs, constructing a classification system, and comparing their strengths and weaknesses. Building on this foundation, this paper focuses on current VAD methods based on MLLMs/LLMs. Finally, based on the trajectory of technological advancements and existing bottlenecks, this paper distills key challenges and outlines future research directions, offering guidance for the VAD community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。