arXiv:2603.26745cs.CV2026-03中稿 · IEEE ICME 2026

用分层语义建模提升隐私保护视频异常检测效果

Motion Semantics Guided Normalizing Flow for Privacy-Preserving Video Anomaly Detection

  • 将动作分解为可解释的语义单元,分层建模运动特征
  • 在HR-ShanghaiTech和HR-UBnormal上分别达到88.1%和75.8%的AUC
  • 适合需要保护隐私又需精准识别异常行为的场景

随着具身感知系统在交互式多媒体应用中日益普及,保护物理环境中人类活动的隐私变得至关重要。视频异常检测是此类系统中智能监控与司法分析的关键任务。基于骨架的方法通过抽象人体姿态表示处理现实世界信息,同时丢弃身份、面部等敏感视觉特征,成为一种隐私保护替代方案。然而,现有骨架方法大多以整体方式建模连续运动轨迹,未能捕捉由离散语义单元和细粒度运动细节构成的人类行为层次结构,导致异常在不同抽象层级下难以区分。为此,我们提出运动语义引导的归一化流(MSG-Flow),将骨架化视频异常检测分解为分层运动语义建模。该方法采用向量量化变分自编码器将连续运动离散化为可解释的语义单元,使用自回归Transformer建模语义级时序依赖,并借助条件归一化流捕捉细节级姿态变化。在HR-ShanghaiTech与HR-UBnormal基准上的大量实验表明,MSG-Flow分别取得88.1%和75.8%的AUC,达到当前最优性能。

原文摘要 · Abstract (English)

As embodied perception systems increasingly bridge digital and physical realms in interactive multimedia applications, the need for privacy-preserving approaches to understand human activities in physical environments has become paramount. Video anomaly detection is a critical task in such embodied multimedia systems for intelligent surveillance and forensic analysis. Skeleton-based approaches have emerged as a privacy-preserving alternative that processes physical world information through abstract human pose representations while discarding sensitive visual attributes such as identity and facial features. However, existing skeleton-based methods predominantly model continuous motion trajectories in a monolithic manner, failing to capture the hierarchical nature of human activities composed of discrete semantic primitives and fine-grained kinematic details, which leads to reduced discriminability when anomalies manifest at different abstraction levels. In this regard, we propose Motion Semantics Guided Normalizing Flow (MSG-Flow) that decomposes skeleton-based VAD into hierarchical motion semantics modeling. It employs vector quantized variational auto-encoder to discretize continuous motion into interpretable primitives, an autoregressive Transformer to model semantic-level temporal dependencies, and a conditional normalizing flow to capture detail-level pose variations. Extensive experiments on benchmarks (HR-ShanghaiTech & HR-UBnormal) demonstrate that MSG-Flow achieves state-of-the-art performance with 88.1% and 75.8% AUC respectively.

视频异常检测隐私保护分层建模归一化流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。