arXiv:2511.05564cs.CV2025-11被引 1

用Mamba设计多尺度时空模型,提升视频异常检测精度与效率

M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly Detection

  • 基于Mamba构建多粒度时空编码器,捕捉不同尺度的运动动态
  • 在三个数据集上分别达到98.5%、92.1%、77.9%的帧级AUC
  • 仅需20.1G FLOPs和45 FPS,适合实际监控系统部署

视频异常检测(VAD)在图像处理领域具有重要应用前景,尤其在视频监控中面临检测精度与计算效率之间的平衡挑战。随着视频内容日益复杂,传统方法在建模多样行为模式和上下文场景时表现不足,且普遍存在时空建模不完整或计算开销过大的问题。本文提出一种基于Mamba的多尺度时空学习框架(M2S2L),通过分层空间编码器在多粒度上建模空间特征,并利用多时间编码器捕获不同时间尺度的运动动态。同时引入特征分解机制,实现外观与运动重建的任务特异性优化,从而支持更精细的行为建模和质量感知的异常评估。在UCSD Ped2、CUHK Avenue和ShanghaiTech三个基准数据集上的实验表明,M2S2L分别取得98.5%、92.1%和77.9%的帧级AUC,且仅需20.1G FLOPs和45 FPS的推理速度,具备实际部署可行性。

原文摘要 · Abstract (English)

Video anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content becomes increasingly complex with diverse behavioral patterns and contextual scenarios, traditional VAD approaches struggle to provide robust assessment for modern surveillance systems. Existing methods either lack comprehensive spatial-temporal modeling or require excessive computational resources for real-time applications. In this regard, we present a Mamba-based multi-scale spatial-temporal learning (M2S2L) framework in this paper. The proposed method employs hierarchical spatial encoders operating at multiple granularities and multi-temporal encoders capturing motion dynamics across different time scales. We also introduce a feature decomposition mechanism to enable task-specific optimization for appearance and motion reconstruction, facilitating more nuanced behavioral modeling and quality-aware anomaly assessment. Experiments on three benchmark datasets demonstrate that M2S2L framework achieves 98.5%, 92.1%, and 77.9% frame-level AUCs on UCSD Ped2, CUHK Avenue, and ShanghaiTech respectively, while maintaining efficiency with 20.1G FLOPs and 45 FPS inference speed, making it suitable for practical surveillance deployment.

视频异常检测Mamba多尺度建模实时部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。