arXiv:2501.01770cs.CV2025-01AAAI被引 45

用隐式姿态代理捕捉动作序列的复杂时序关系,提升3D人体姿态估计精度。

TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose Estimation

  • 引入隐式姿态代理作为中间表示,分步建模时序相关性。
  • 在Human3.6M和MPI-INF-3DHP上优于现有最先进方法。
  • 适合关注多帧人体姿态时序建模的研究者。

近期的多帧提升方法主导了3D人体姿态估计。然而,先前方法忽略了2D姿态序列中的复杂依赖关系,仅学习单一的时序相关性。为缓解这一限制,我们提出TCPFormer,利用隐式姿态代理作为中间表示。每个隐式姿态代理可建立一种时序相关性,从而帮助我们学习更全面的人体运动时序特征。具体而言,该方法包含三个关键组件:代理更新模块(PUM)、代理调用模块(PIM)和代理注意力模块(PAM)。PUM首先使用姿态特征更新隐式姿态代理,使其存储来自姿态序列的代表性信息;PIM随后调用并融合姿态代理与姿态序列,增强每个姿态的运动语义;最后,PAM利用姿态序列与姿态代理之间的映射关系,强化整个姿态序列的时序相关性。在Human3.6M和MPI-INF-3DHP数据集上的实验表明,所提出的TCPFormer优于之前的最先进方法。

原文摘要 · Abstract (English)

Recent multi-frame lifting methods have dominated the 3D human pose estimation. However, previous methods ignore the intricate dependence within the 2D pose sequence and learn single temporal correlation. To alleviate this limitation, we propose TCPFormer, which leverages an implicit pose proxy as an intermediate representation. Each proxy within the implicit pose proxy can build one temporal correlation therefore helping us learn more comprehensive temporal correlation of human motion. Specifically, our method consists of three key components: Proxy Update Module (PUM), Proxy Invocation Module (PIM), and Proxy Attention Module (PAM). PUM first uses pose features to update the implicit pose proxy, enabling it to store representative information from the pose sequence. PIM then invocates and integrates the pose proxy with the pose sequence to enhance the motion semantics of each pose. Finally, PAM leverages the above mapping between the pose sequence and pose proxy to enhance the temporal correlation of the whole pose sequence. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that our proposed TCPFormer outperforms the previous state-of-the-art methods.

3D姿态估计时序建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。