arXiv:2412.09026cs.CV2024-12AAAI被引 34

用局部扩散模型联合建模运动与外观,提升小目标异常检测精度

Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model

  • 基于补丁的扩散模型,专注捕捉细粒度局部信息
  • 在4个数据集上均超越现有方法,异常检测更精准
  • 适合关注视频监控中细微异常行为的科研与工程人员

近期视频异常检测研究将任务视为生成问题,训练扩散模型仅恢复正常模式,从而将异常视为离群点。然而,现有方法忽略异常形态多样性,且在特征层面预测正常样本,而监控视频中的异常对象通常较小。为此,提出一种新型基于补丁的扩散模型,专门捕捉细粒度局部信息。进一步观察到异常在视频中表现为外观与运动的双重偏离,因此主张需同时考虑二者以实现准确帧预测。为此,引入创新的运动与外观条件,并无缝集成至补丁扩散模型中,引导模型生成语义内容与运动关系一致的合理预测。在四个挑战性视频异常检测数据集上的实验结果验证了该方法的有效性,表明其始终优于大多数现有方法。

原文摘要 · Abstract (English)

A recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recover normal patterns exclusively, thus reporting abnormal patterns as outliers. Yet, existing attempts neglect the various formations of anomaly and predict normal samples at the feature level regardless that abnormal objects in surveillance videos are often relatively small. To address this, a novel patch-based diffusion model is proposed, specifically engineered to capture fine-grained local information. We further observe that anomalies in videos manifest themselves as deviations in both appearance and motion. Therefore, we argue that a comprehensive solution must consider both of these aspects simultaneously to achieve accurate frame prediction. To address this, we introduce innovative motion and appearance conditions that are seamlessly integrated into our patch diffusion model. These conditions are designed to guide the model in generating coherent and contextually appropriate predictions for both semantic content and motion relations. Experimental results in four challenging video anomaly detection datasets empirically substantiate the efficacy of our proposed approach, demonstrating that it consistently outperforms most existing methods in detecting abnormal behaviors.

视频异常检测扩散模型运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。