通过语义动态与时空协同,提升视频人体姿态估计精度。
Learning semantical dynamics and spatiotemporal collaboration for human pose estimation in video
- 设计多粒度语义运动编码器,利用掩码重建学习帧间语义关联。
- 在三个数据集上达到新最好性能,PoseTrack2017 AP达64.5。
- 适合需要高精度视频姿态估计的场景,如动作识别与追踪。
时序建模与时空协同是视频人体姿态估计的关键技术。现有方法多依赖光流或时序差分,在像素层面学习局部视觉对应关系以捕捉运动动态,但这类方法本质上依赖像素级相似性,忽视帧间语义关联,且易受遮挡、模糊等图像退化影响。此外,多数方法通过简单拼接或相加融合运动与空间特征,难以充分挖掘两种模态的潜力。本文提出一种新框架,学习多层级语义动态与密集时空协同,用于多帧人体姿态估计。首先设计多层级语义运动编码器,采用多掩码上下文与姿态重建策略,通过逐步掩码(块)立方体和帧特征,激发模型探索多粒度时空语义关系。进一步引入空间-运动互学模块,密集传播并整合空间与运动特征中的上下文信息,提升模型表达能力。大量实验表明,该方法在三个基准数据集上取得新最优结果:PoseTrack2017 AP 64.5,PoseTrack2018 AP 65.3,PoseTrack21 AP 63.9。
原文摘要 · Abstract (English)
Temporal modeling and spatio-temporal collaboration are pivotal techniques for video-based human pose estimation. Most state-of-the-art methods adopt optical flow or temporal difference, learning local visual content correspondence across frames at the pixel level, to capture motion dynamics. However, such a paradigm essentially relies on localized pixel-to-pixel similarity, which neglects the semantical correlations among frames and is vulnerable to image quality degradations (e.g. occlusions or blur). Moreover, existing approaches often combine motion and spatial (appearance) features via simple concatenation or summation, leading to practical challenges in fully leveraging these distinct modalities. In this paper, we present a novel framework that learns multi-level semantical dynamics and dense spatio-temporal collaboration for multi-frame human pose estimation. Specifically, we first design a Multi-Level Semantic Motion Encoder using a multi-masked context and pose reconstruction strategy. This strategy stimulates the model to explore multi-granularity spatiotemporal semantic relationships among frames by progressively masking the features of (patch) cubes and frames. We further introduce a Spatial-Motion Mutual Learning module which densely propagates and consolidates context information from spatial and motion features to enhance the capability of the model. Extensive experiments demonstrate that our approach sets new state-of-the-art results on three benchmark datasets, PoseTrack2017, PoseTrack2018, and PoseTrack21.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。