arXiv:2410.04574cs.CVcs.LG2024-10被引 12

用双变压器融合提升遮挡下的3D人体姿态估计精度

Enhancing 3D Human Pose Estimation Amidst Severe Occlusion with Dual Transformer Fusion

  • 通过时间插值生成遮挡引导,补全缺失关节点
  • 双中间视图自优化后融合,实现端到端训练
  • 在Human3.6M和MPI-INF-3DHP上超越现有方法

单目视频中的3D人体姿态估计面临多类型遮挡的严峻挑战。以往研究利用空间与时间线索从2D关节观测中推断3D姿态。本文提出一种双变压器融合(DTF)算法,可在严重遮挡下实现整体3D姿态估计。针对遮挡导致的关节点缺失问题,我们设计基于时间插值的遮挡引导机制。该方法采用创新的DTF架构,先生成一对中间视图,每个中间视图通过自精修方案进行空间优化,随后融合得到最终3D人体姿态估计。整个系统可端到端训练。在Human3.6M和MPI-INF-3DHP数据集上的大量实验表明,该方法显著优于现有最先进方法。

原文摘要 · Abstract (English)

In the field of 3D Human Pose Estimation from monocular videos, the presence of diverse occlusion types presents a formidable challenge. Prior research has made progress by harnessing spatial and temporal cues to infer 3D poses from 2D joint observations. This paper introduces a Dual Transformer Fusion (DTF) algorithm, a novel approach to obtain a holistic 3D pose estimation, even in the presence of severe occlusions. Confronting the issue of occlusion-induced missing joint data, we propose a temporal interpolation-based occlusion guidance mechanism. To enable precise 3D Human Pose Estimation, our approach leverages the innovative DTF architecture, which first generates a pair of intermediate views. Each intermediate-view undergoes spatial refinement through a self-refinement schema. Subsequently, these intermediate-views are fused to yield the final 3D human pose estimation. The entire system is end-to-end trainable. Through extensive experiments conducted on the Human3.6M and MPI-INF-3DHP datasets, our method's performance is rigorously evaluated. Notably, our approach outperforms existing state-of-the-art methods on both datasets, yielding substantial improvements. The code is available here: https://github.com/MehwishG/DTF.

3D姿态估计遮挡处理双变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。