arXiv:2511.10334cs.CV2025-11中稿 · AAAI被引 8

通过解耦语义对齐,提升弱监督视频异常检测的细粒度区分能力。

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

  • 分粗粒度与细粒度两阶段解耦正常与异常特征。
  • 在XD-Violence和UCF-Crime上优于现有方法。
  • 适合需要精准识别异常类别的实际安防场景。

近期弱监督视频异常检测借助多模态基础模型(如CLIP)与多实例学习范式,显著提升了性能。然而,其目标易聚焦于最显著响应片段,忽视多样正常模式的挖掘,且因外观相似导致类别混淆,影响细粒度分类效果。为此,提出解耦语义对齐网络DSANet,从粗粒度与细粒度层面显式分离异常与正常特征,增强可区分性。粗粒度层面引入自引导正常建模分支,基于学习到的正常原型重建输入视频特征,促进模型利用视频中固有的正常线索,提升正常模式与异常事件的时间分离性。细粒度层面设计解耦对比语义对齐机制:先用帧级异常得分将视频分解为事件中心与背景中心成分,再进行视觉-语言对比学习,强化类别判别表征。在两个标准基准XD-Violence与UCF-Crime上的实验表明,DSANet优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advancements in weakly-supervised video anomaly detection have achieved remarkable performance by applying the multiple instance learning paradigm based on multimodal foundation models such as CLIP to highlight anomalous instances and classify categories. However, their objectives may tend to detect the most salient response segments, while neglecting to mine diverse normal patterns separated from anomalies, and are prone to category confusion due to similar appearance, leading to unsatisfactory fine-grained classification results. Therefore, we propose a novel Disentangled Semantic Alignment Network (DSANet) to explicitly separate abnormal and normal features from coarse-grained and fine-grained aspects, enhancing the distinguishability. Specifically, at the coarse-grained level, we introduce a self-guided normality modeling branch that reconstructs input video features under the guidance of learned normal prototypes, encouraging the model to exploit normality cues inherent in the video, thereby improving the temporal separation of normal patterns and anomalous events. At the fine-grained level, we present a decoupled contrastive semantic alignment mechanism, which first temporally decomposes each video into event-centric and background-centric components using frame-level anomaly scores and then applies visual-language contrastive learning to enhance class-discriminative representations. Comprehensive experiments on two standard benchmarks, namely XD-Violence and UCF-Crime, demonstrate that DSANet outperforms existing state-of-the-art methods.

视频异常检测弱监督解耦学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。