用双条件扩散模型提升姿态异常检测,兼顾结构与动态异常。
Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly Detection
- 引入双条件扩散框架,融合姿态特征与语义信息
- 在四个数据集上显著超越现有方法,最优指标达85.3%
- 适合需要高精度动作异常识别的智能监控场景
视频异常检测(VAD)是计算机视觉的重要任务。现有方法分为基于重建和基于预测两类:前者擅长发现不规则结构,后者能捕捉异常行为趋势。本文针对姿态驱动的视频异常检测,提出一种新框架——双条件运动扩散(DCMD),融合两类方法优势。该框架通过条件化运动和条件化嵌入,分别利用人体动作的姿态特性与潜在语义。在反向扩散过程中,设计运动变换器,从人体运动频谱空间的多层特征中捕获潜在关联。为增强正常与异常样本的区分能力,提出新型统一关联差异(UAD)正则化,基于高斯核时间关联与自注意力全局关联。推理阶段引入掩码补全策略,提升条件运动对异常检测预测分支的利用效率。在四个数据集上的大量实验表明,本方法显著优于现有最优方法,展现出更强泛化能力。
原文摘要 · Abstract (English)
Video Anomaly Detection (VAD) is essential for computer vision research. Existing VAD methods utilize either reconstruction-based or prediction-based frameworks. The former excels at detecting irregular patterns or structures, whereas the latter is capable of spotting abnormal deviations or trends. We address pose-based video anomaly detection and introduce a novel framework called Dual Conditioned Motion Diffusion (DCMD), which enjoys the advantages of both approaches. The DCMD integrates conditioned motion and conditioned embedding to comprehensively utilize the pose characteristics and latent semantics of observed movements, respectively. In the reverse diffusion process, a motion transformer is proposed to capture potential correlations from multi-layered characteristics within the spectrum space of human motion. To enhance the discriminability between normal and abnormal instances, we design a novel United Association Discrepancy (UAD) regularization that primarily relies on a Gaussian kernel-based time association and a self-attention-based global association. Finally, a mask completion strategy is introduced during the inference stage of the reverse diffusion process to enhance the utilization of conditioned motion for the prediction branch of anomaly detection. Extensive experiments on four datasets demonstrate that our method dramatically outperforms state-of-the-art methods and exhibits superior generalization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。