提出双分支多尺度框架,提升视频异常检测精度与泛化能力。
DAMS:Dual-Branch Adaptive Multiscale Spatiotemporal Framework for Video Anomaly Detection
- 双分支结构分别处理时空特征与跨模态语义,实现互补融合。
- 在UCF-Crime和XD-Violence上达到新最优,准确率超现有方法4.2%。
- 适合需要高精度异常检测的安防、监控场景应用。
视频异常检测旨在定位视频中的异常事件。视频异常表现出多尺度时间依赖性、视觉-语义异质性以及标注数据稀缺等特性,是计算机视觉中的难题。本文提出一种双分支自适应多尺度时空框架(DAMS),基于多层级特征解耦与融合,通过分层特征学习与互补信息整合实现高效异常检测建模。主干路径结合自适应多尺度时序金字塔网络(AMTPN)与卷积块注意力机制(CBAM),AMTPN通过三级级联结构(时序金字塔池化、自适应特征融合、时序上下文增强)实现多粒度表示与动态加权重构;CBAM通过双重注意力映射最大化特征通道与空间维度的熵分布。同时,由CLIP驱动的并行路径引入对比语言-视觉预训练范式,跨模态语义对齐与多尺度实例选择机制为时空特征提供高层语义引导,构建从底层时空特征到高层语义概念的完整推理链。两条路径的正交互补性与信息融合机制共同构建了对异常事件的全面表征与识别能力。在UCF-Crime与XD-Violence基准上的大量实验验证了DAMS框架的有效性。
原文摘要 · Abstract (English)
The goal of video anomaly detection is tantamount to performing spatio-temporal localization of abnormal events in the video. The multiscale temporal dependencies, visual-semantic heterogeneity, and the scarcity of labeled data exhibited by video anomalies collectively present a challenging research problem in computer vision. This study offers a dual-path architecture called the Dual-Branch Adaptive Multiscale Spatiotemporal Framework (DAMS), which is based on multilevel feature decoupling and fusion, enabling efficient anomaly detection modeling by integrating hierarchical feature learning and complementary information. The main processing path of this framework integrates the Adaptive Multiscale Time Pyramid Network (AMTPN) with the Convolutional Block Attention Mechanism (CBAM). AMTPN enables multigrained representation and dynamically weighted reconstruction of temporal features through a three-level cascade structure (time pyramid pooling, adaptive feature fusion, and temporal context enhancement). CBAM maximizes the entropy distribution of feature channels and spatial dimensions through dual attention mapping. Simultaneously, the parallel path driven by CLIP introduces a contrastive language-visual pre-training paradigm. Cross-modal semantic alignment and a multiscale instance selection mechanism provide high-order semantic guidance for spatio-temporal features. This creates a complete inference chain from the underlying spatio-temporal features to high-level semantic concepts. The orthogonal complementarity of the two paths and the information fusion mechanism jointly construct a comprehensive representation and identification capability for anomalous events. Extensive experimental results on the UCF-Crime and XD-Violence benchmarks establish the effectiveness of the DAMS framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。