无需标注数据,通过伪标签引导全局变换实现跨域3D人体姿态估计
Unsupervised Cross-Domain 3D Human Pose Estimation via Pseudo-Label-Guided Global Transforms
- 用伪标签生成+全局变换对齐不同视角下的姿态位置
- 在多个跨数据集测试中超越当前最佳方法,接近有监督模型表现
- 适合无目标域标注、需跨场景迁移的3D姿态估计任务
现有3D人体姿态估计方法在跨场景推理时性能下降,主要源于相机视角、位置、姿势和体型等特征差异。其中,相机视角与位置显著影响姿态的全局空间分布。为此,本文提出一种新框架,显式对源域与目标域的相机坐标系中姿态位置进行全局变换。首先通过伪标签生成模块,基于目标数据集的2D姿态生成伪3D姿态;再利用以人体为中心的坐标系作为桥梁,实现跨域姿态位置方向的一致对齐,确保空间参照统一;为进一步提升泛化能力,引入姿态增强模块应对姿势与体型变化。该过程迭代进行,使优化后的伪标签持续改进域适应效果。在Human3.6M、MPI-INF-3DHP和3DPW等多个跨数据集基准上评估,所提方法优于当前最优方法,甚至超过在目标域上训练的监督模型。
原文摘要 · Abstract (English)
Existing 3D human pose estimation methods often suffer in performance, when applied to cross-scenario inference, due to domain shifts in characteristics such as camera viewpoint, position, posture, and body size. Among these factors, camera viewpoints and locations have been shown to contribute significantly to the domain gap by influencing the global positions of human poses. To address this, we propose a novel framework that explicitly conducts global transformations between pose positions in the camera coordinate systems of source and target domains. We start with a Pseudo-Label Generation Module that is applied to the 2D poses of the target dataset to generate pseudo-3D poses. Then, a Global Transformation Module leverages a human-centered coordinate system as a novel bridging mechanism to seamlessly align the positional orientations of poses across disparate domains, ensuring consistent spatial referencing. To further enhance generalization, a Pose Augmentor is incorporated to address variations in human posture and body size. This process is iterative, allowing refined pseudo-labels to progressively improve guidance for domain adaptation. Our method is evaluated on various cross-dataset benchmarks, including Human3.6M, MPI-INF-3DHP, and 3DPW. The proposed method outperforms state-of-the-art approaches and even outperforms the target-trained model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。