用点扩散模型提升个性化3D人体姿态估计精度
PHD: Personalized 3D Human Body Fitting with Point Diffusion
- 先校准用户体型,再用体型条件化扩散模型优化姿态
- 在Pelvis对齐和绝对姿态误差上均优于现有方法
- 仅需合成数据训练,可无缝接入现有姿态估计算法
我们提出PHD,一种基于用户特定体型信息的个性化3D人体网格恢复与体态拟合新方法,以提升视频中姿态估计的准确性。传统方法为通用设计,依赖2D图像约束优化姿态,但忽视个体体型差异,导致3D精度下降。本文通过解耦流程:先校准用户体型,再基于该体型条件进行姿态拟合。核心是构建体形条件化的3D姿态先验,采用点扩散变换器实现,并通过点蒸馏采样损失迭代引导拟合过程,有效缓解对2D约束的过度依赖。实验表明,该方法显著提升骨盆对齐姿态精度及绝对姿态精度(关键指标常被忽略)。此外,方法仅需合成数据训练,具备高数据效率,且可作为即插即用模块,无缝集成至现有3D姿态估计器中提升性能。
原文摘要 · Abstract (English)
We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。