系统梳理深度学习时代视频异常检测的各类方法与进展
Deep Learning for Video Anomaly Detection: A Review
- 按监督程度分类,涵盖五类主流VAD方法
- 对比分析不同方法性能,覆盖全监督到开放集场景
- 整合数据集、代码与评测标准,适合研究者快速入门
视频异常检测(VAD)旨在发现视频中偏离正常模式的行为或事件。作为计算机视觉领域长期任务,近年来在深度学习推动下取得显著进展。随着模型架构不断演进,大量基于深度学习的方法涌现,显著提升了算法泛化能力并拓展了应用范围。然而方法繁多、文献庞杂,亟需全面综述。本文系统回顾了五类VAD方法:半监督、弱监督、全监督、无监督及开放集监督,并深入探讨基于预训练大模型的最新进展,弥补以往综述仅关注半监督和小模型的局限。针对不同监督水平,构建清晰分类体系,深入分析方法特性并比较性能表现。同时涵盖所有相关公共数据集、开源代码与评估指标。最后,提出若干关键研究方向,为该领域发展提供指引。
原文摘要 · Abstract (English)
Video anomaly detection (VAD) aims to discover behaviors or events deviating from the normality in videos. As a long-standing task in the field of computer vision, VAD has witnessed much good progress. In the era of deep learning, with the explosion of architectures of continuously growing capability and capacity, a great variety of deep learning based methods are constantly emerging for the VAD task, greatly improving the generalization ability of detection algorithms and broadening the application scenarios. Therefore, such a multitude of methods and a large body of literature make a comprehensive survey a pressing necessity. In this paper, we present an extensive and comprehensive research review, covering the spectrum of five different categories, namely, semi-supervised, weakly supervised, fully supervised, unsupervised and open-set supervised VAD, and we also delve into the latest VAD works based on pre-trained large models, remedying the limitations of past reviews in terms of only focusing on semi-supervised VAD and small model based methods. For the VAD task with different levels of supervision, we construct a well-organized taxonomy, profoundly discuss the characteristics of different types of methods, and show their performance comparisons. In addition, this review involves the public datasets, open-source codes, and evaluation metrics covering all the aforementioned VAD tasks. Finally, we provide several important research directions for the VAD community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。