用动作典型性与上下文独特性提升骨骼数据异常检测泛化能力
Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
- 通过语言模型学习动作语义典型性,捕捉正常与异常行为模式
- 测试时分析时空差异,自适应构建场景边界,无需目标域训练数据
- 在4个大规模数据集上超越现有方法,适用于100多个未见监控场景
零样本视频异常检测(ZS-VAD)需在无目标域训练数据情况下定位异常,对隐私保护和新监控部署至关重要。基于骨骼的方法因消除背景与人体外观差异,具有天然泛化优势。然而现有方法仅学习低层骨骼表征,依赖受限于领域的正常性边界,难以适应新场景中的不同正常与异常行为模式。本文提出一种新型零样本视频异常检测框架,通过动作典型性与唯一性学习挖掘骨骼数据潜力。首先引入语言引导的语义典型性建模模块,将骨骼片段映射至动作语义空间,并在训练中提炼大语言模型对典型正常与异常行为的知识。其次提出测试时上下文唯一性分析模块,精细分析骨骼片段间的时空差异,从而推导出场景自适应的边界。无需使用目标域任何训练样本,本方法在四个大规模VAD数据集ShanghaiTech、UBnormal、NWPU和UCF-Crime上达到当前最优表现,涵盖超过100个未见监控场景。
原文摘要 · Abstract (English)
Zero-Shot Video Anomaly Detection (ZS-VAD) requires temporally localizing anomalies without target domain training data, which is a crucial task due to various practical concerns, e.g., data privacy or new surveillance deployments. Skeleton-based approach has inherent generalizable advantages in achieving ZS-VAD as it eliminates domain disparities both in background and human appearance. However, existing methods only learn low-level skeleton representation and rely on the domain-limited normality boundary, which cannot generalize well to new scenes with different normal and abnormal behavior patterns. In this paper, we propose a novel zero-shot video anomaly detection framework, unlocking the potential of skeleton data via action typicality and uniqueness learning. Firstly, we introduce a language-guided semantic typicality modeling module that projects skeleton snippets into action semantic space and distills LLM's knowledge of typical normal and abnormal behaviors during training. Secondly, we propose a test-time context uniqueness analysis module to finely analyze the spatio-temporal differences between skeleton snippets and then derive scene-adaptive boundaries. Without using any training samples from the target domain, our method achieves state-of-the-art results against skeleton-based methods on four large-scale VAD datasets: ShanghaiTech, UBnormal, NWPU, and UCF-Crime, featuring over 100 unseen surveillance scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。