arXiv:2511.13204cs.CV2025-11AAAI被引 2

通过语义引导特征重校准,提升弱监督视频异常检测精度

RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection

  • 引入运动感知时序注意力与类别导向精炼模块,双路径建模异常
  • 在WVAD数据集上达到新最优,异常检测准确率显著提升
  • 适合需要细粒度异常识别的工业监控场景

弱监督视频异常检测仅依赖视频级标签,兼顾标注效率与实际应用。然而现有方法常将所有异常事件视为单一类别,忽略了真实世界异常在语义与时间上的多样性。受人类感知异常方式启发,我们提出RefineVAD框架,模拟对时序运动模式与语义结构的联合理解。该框架包含两个核心模块:一是运动感知时序注意力与重校准(MoTAR),通过基于位移的注意力和全局Transformer建模动态调整时间关注;二是类别导向精炼(CORE),利用交叉注意力将片段级特征与可学习类别原型对齐,注入软异常类别先验。通过联合建模时序动态与语义结构,显式捕捉“运动如何演变”及“对应何种语义类别”。在WVAD基准上的大量实验验证了RefineVAD的有效性,并凸显融合语义上下文对引导特征精炼至异常相关模式的重要性。

原文摘要 · Abstract (English)

Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by treating all abnormal events as a single category, overlooking the diverse semantic and temporal characteristics intrinsic to real-world anomalies. Inspired by how humans perceive anomalies, by jointly interpreting temporal motion patterns and semantic structures underlying different anomaly types, we propose RefineVAD, a novel framework that mimics this dual-process reasoning. Our framework integrates two core modules. The first, Motion-aware Temporal Attention and Recalibration (MoTAR), estimates motion salience and dynamically adjusts temporal focus via shift-based attention and global Transformer-based modeling. The second, Category-Oriented Refinement (CORE), injects soft anomaly category priors into the representation space by aligning segment-level features with learnable category prototypes through cross-attention. By jointly leveraging temporal dynamics and semantic structure, explicitly models both "how" motion evolves and "what" semantic category it resembles. Extensive experiments on WVAD benchmark validate the effectiveness of RefineVAD and highlight the importance of integrating semantic context to guide feature refinement toward anomaly-relevant patterns.

视频异常检测弱监督学习语义引导特征重校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。