arXiv:2409.03236cs.CV2024-09被引 18

通过解耦场景与动作,提升复杂场景下人体异常检测准确率

Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly Detection

  • 将场景与动作分离建模,再交叉融合以增强理解
  • 在UCF-Crime、HMDB51等数据集上显著优于现有方法
  • 适用于有监督、弱监督和无监督多种训练场景

人体相关视频中的异常检测对监控应用至关重要。现有方法主要分为基于外观和基于动作两类。外观方法依赖颜色、纹理等低级视觉特征,在熟悉场景中表现良好,但在未知或变化场景下泛化能力差;动作方法关注行为本身,却忽略场景信息,导致误判(如海滩跑步与街面跑步均被视为正常)。二者难以融合,限制了复杂场景下的检测效果。为此,本文提出一种基于解耦的异常检测架构DecoAD,通过解耦并重新交织场景与动作特征,实现对复杂行为与环境的更精准理解。该模型支持全监督、弱监督和无监督设置,在UCF-Crime、HMDB51等多个数据集上取得显著性能提升。

原文摘要 · Abstract (English)

Detecting anomalies in human-related videos is crucial for surveillance applications. Current methods primarily include appearance-based and action-based techniques. Appearance-based methods rely on low-level visual features such as color, texture, and shape. They learn a large number of pixel patterns and features related to known scenes during training, making them effective in detecting anomalies within these familiar contexts. However, when encountering new or significantly changed scenes, i.e., unknown scenes, they often fail because existing SOTA methods do not effectively capture the relationship between actions and their surrounding scenes, resulting in low generalization. In contrast, action-based methods focus on detecting anomalies in human actions but are usually less informative because they tend to overlook the relationship between actions and their scenes, leading to incorrect detection. For instance, the normal event of running on the beach and the abnormal event of running on the street might both be considered normal due to the lack of scene information. In short, current methods struggle to integrate low-level visual and high-level action features, leading to poor anomaly detection in varied and complex scenes. To address this challenge, we propose a novel decoupling-based architecture for human-related video anomaly detection (DecoAD). DecoAD significantly improves the integration of visual and action features through the decoupling and interweaving of scenes and actions, thereby enabling a more intuitive and accurate understanding of complex behaviors and scenes. DecoAD supports fully supervised, weakly supervised, and unsupervised settings.

视频异常检测知识图谱解耦建模场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。