提升动作检测对视角变化和时间一致性的鲁棒性
Improving Viewpoint-Invariance and Temporal Consistency for Action Detection

- 通过虚拟视角增强训练,提升模型视角不变性
- 多尺度时序编码器在PKU-MMD和BABEL上超越当前最佳
- 适合需要高精度动作识别的视频分析场景
视角变化不变性和动作时间一致性是未剪辑视频中人体动作检测有效部署的关键。现有基于外观的视频检测方法在训练时视角多样性有限,而基于运动的方法难以建模连续运动窗口间的细粒度时间关系。本文提出一种两阶段动作检测方法,第一阶段从增强的虚拟视角中提取运动特征,仅用于训练;第二阶段引入基于选择性状态空间序列建模的新颖视角不变、多尺度时序编码器,以跨视角和时间尺度聚合信息。在PKU-MMD和BABEL基准上的实验表明,该方法在所有评估划分中均显著优于现有最先进方法。代码与训练模型见:https://icb-vision-ai.github.io/HydraView-TAD
原文摘要 · Abstract (English)
Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed videos. Existing appearance-based video detection methods often struggle with limited viewpoint diversity during training, while motion-based detection approaches frequently fail to model fine-grained temporal relationships across consecutive motion windows. This paper introduces a novel two-stage action detection approach designed to improve both view-invariance and global temporal coherence properties. In the first stage, we extract motion features from augmented virtual viewpoints, solely used at training. Then, the second stage introduces a new view-invariant, multi-scale temporal encoder based on selective state-space sequence modelling to aggregate information across viewpoints and time scales. Experiments on PKU-MMD and BABEL benchmarks demonstrate that this approach significantly outperforms state-of-the-art methods in all considered splits. Code and trained models are available at: https://icb-vision-ai.github.io/HydraView-TAD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。