用主成分分析提升动作异常检测的稳定性与精度
STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection

- 将骨骼序列投影到主成分空间,避免噪声导致的不合理姿态
- 在UBnormal数据集上达90.1%的AUROC,比之前方法高12.2%
- 适合需要实时处理且对姿态估计误差敏感的应用场景
基于骨骼的视频异常检测(VAD)提供了一种鲁棒且保护隐私的异常行为识别方案。为建模正常静态与运动姿态分布,现有方法通过去噪得分匹配(DSM)训练能量模型(EBMs)。然而,直接向原始关节坐标注入噪声会生成物理上不可能的姿态,且随时间窗口扩大,结构崩溃问题愈发严重。为此,我们提出STEP框架,利用主成分分析(PCA)将姿态序列投影至紧凑、白化后的主成分空间。在该结构良好的空间中学习数据密度,使注入噪声转化为物理合理的变异,从而实现更长视频序列的稳定处理,避免原始坐标基线的性能下降。此外,为缓解遮挡或运动模糊带来的姿态估计误差,我们引入基于置信度分数的序列级加权机制。本框架计算效率高、结构轻量,在挑战性数据集UBnormal上达到90.1% AUROC,较此前骨架类最先进方法提升12.2%,并在ShanghaiTech基准上取得具有竞争力的结果。
原文摘要 · Abstract (English)
Skeleton-based Video Anomaly Detection (VAD) offers a robust, privacy-preserving solution for identifying abnormal behaviors. To model the distribution of normal static and moving poses, recent methods train Energy-Based Models (EBMs) via Denoising Score Matching (DSM). However, directly injecting noise, required for training, into raw joint coordinates creates physically impossible poses, and this structural collapse severely worsens as the temporal window expands. To address this, we introduce STEP, a simple framework that utilizes Principal Component Analysis (PCA) to project pose sequences into a compact, whitened PC-space. Learning the data density within this well-behaved PC-space ensures that the injected noise translates into physically plausible variations, which allows the model to process longer video sequences without the performance collapse of raw coordinate baselines. Additionally, to mitigate inherent pose estimation inaccuracies arising from occlusions or motion blur, we integrate a sequence-level weighting mechanism based on the estimator's confidence scores. Operating at real-time computational efficiency, our simple and lightweight framework outperforms the previous skeleton-based state-of-the-art by 12.2% (90.1% AUROC) on the challenging UBnormal dataset and achieves highly competitive results by improving on the ShanghaiTech benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。