通过因果一致性学习,让模型在无监督下更稳定地识别监控视频中的异常行为。
CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos
- 基于因果模型,分离场景干扰,提取鲁棒的正常模式表征
- 在多个数据集上优于传统方法,尤其在跨场景、数据少时表现更稳
- 适合需要抗干扰、少标注的实时监控异常检测场景
视频异常检测(VAD)在信息取证与公共安全等领域具有重要应用价值。由于异常事件稀少且多样,现有方法通常仅依赖易获取的正常事件,在无监督条件下建模时空正常模式。然而,已有研究表明,这些方法在真实场景中无法应对无标签的数据偏移(如场景切换),且因深度网络过度泛化,难以捕捉细微异常。受因果学习启发,本文认为存在能充分泛化正常模式并显著偏离的因果因素。为此提出因果表示一致性学习(CRCL),隐式挖掘无监督视频正常性学习中的场景鲁棒因果变量。具体而言,基于结构因果模型,设计场景去偏学习与因果启发的正常性学习,分别剥离深层表征中的场景混淆因素,学习因果视频正常性。大量实验证明,该方法在基准数据集上优于传统深度表示学习。消融实验与扩展验证表明,CRCL可有效应对多场景下的无标签偏移,且在训练数据有限时仍保持稳定性能。
原文摘要 · Abstract (English)
Video Anomaly Detection (VAD) remains a fundamental yet formidable task in the video understanding community, with promising applications in areas such as information forensics and public safety protection. Due to the rarity and diversity of anomalies, existing methods only use easily collected regular events to model the inherent normality of normal spatial-temporal patterns in an unsupervised manner. Previous studies have shown that existing unsupervised VAD models are incapable of label-independent data offsets (e.g., scene changes) in real-world scenarios and may fail to respond to light anomalies due to the overgeneralization of deep neural networks. Inspired by causality learning, we argue that there exist causal factors that can adequately generalize the prototypical patterns of regular events and present significant deviations when anomalous instances occur. In this regard, we propose Causal Representation Consistency Learning (CRCL) to implicitly mine potential scene-robust causal variable in unsupervised video normality learning. Specifically, building on the structural causal models, we propose scene-debiasing learning and causality-inspired normality learning to strip away entangled scene bias in deep representations and learn causal video normality, respectively. Extensive experiments on benchmarks validate the superiority of our method over conventional deep representation learning. Moreover, ablation studies and extension validation show that the CRCL can cope with label-independent biases in multi-scene settings and maintain stable performance with only limited training data available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。