arXiv:2502.10724cs.CV2025-02ICML被引 1

用视频语义引导3D人体姿态估计,解决测试时适应中的模糊和误判问题。

Semantics-aware Test-time Adaptation for 3D Human Pose Estimation

  • 引入语义感知运动先验,结合视频语义动态调整预测。
  • 在3DPW和3DHP上将PA-MPJPE降低超12%。
  • 适合处理遮挡、截断场景,对动作理解要求高的应用。

本文揭示了3D人体姿态估计中存在语义错位问题。在测试时自适应(TTA)任务中,该问题表现为预测结果过度平滑且缺乏引导,导致模型趋于平均姿态;当出现遮挡或截断时,适应过程完全失去方向。为此,我们首次提出在测试时自适应中融入语义感知的运动先验。通过利用视频理解能力及结构化的运动-文本空间,在测试阶段使模型的姿态预测与视频语义保持一致。此外,我们基于运动-文本相似性完成缺失的2D姿态补全,强化了先验对遮挡和截断情况下的指导作用。实验表明,本方法显著优于现有最优的TTA技术,在3DPW和3DHP数据集上,PA-MPJPE指标下降超过12%。

原文摘要 · Abstract (English)

This work highlights a semantics misalignment in 3D human pose estimation. For the task of test-time adaptation, the misalignment manifests as overly smoothed and unguided predictions. The smoothing settles predictions towards some average pose. Furthermore, when there are occlusions or truncations, the adaptation becomes fully unguided. To this end, we pioneer the integration of a semantics-aware motion prior for the test-time adaptation of 3D pose estimation. We leverage video understanding and a well-structured motion-text space to adapt the model motion prediction to adhere to video semantics during test time. Additionally, we incorporate a missing 2D pose completion based on the motion-text similarity. The pose completion strengthens the motion prior's guidance for occlusions and truncations. Our method significantly improves state-of-the-art 3D human pose estimation TTA techniques, with more than 12% decrease in PA-MPJPE on 3DPW and 3DHP.

姿态估计测试自适应视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。