用扩散模型建模全身人体姿态,提升动作生成鲁棒性。
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior
- 将姿态任务统一为逆问题,通过变分扩散采样求解。
- 在多类姿态数据集上优于现有方法,实现全身体态建模新基准。
- 适合需要高精度全身姿态生成的研究与应用,如动画、虚拟人。
我们提出DPoser-X,一种基于扩散模型的3D全身人体姿态先验模型。由于人体关节结构复杂且高质量全身姿态数据稀缺,构建通用且鲁棒的全身姿态先验仍具挑战。为此,我们引入基于扩散模型的姿态先验(DPoser),并扩展为DPoser-X以支持更丰富的全身姿态建模。该方法将各类姿态相关任务统一为逆问题,通过变分扩散采样求解。为提升下游任务表现,我们设计了一种针对姿态数据特性的截断时间步调度策略,并提出掩码训练机制,有效融合全身与局部数据集,捕捉肢体间依赖关系同时避免特定动作过拟合。大量实验表明,DPoser-X在人体、手部、面部及全身姿态建模多个基准上均表现优异,持续超越当前最优方法,确立了全身体态先验建模的新标准。
原文摘要 · Abstract (English)
We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limitations, we introduce a Diffusion model as body Pose prior (DPoser) and extend it to DPoser-X for expressive whole-body human pose modeling. Our approach unifies various pose-centric tasks as inverse problems, solving them through variational diffusion sampling. To enhance performance on downstream applications, we introduce a novel truncated timestep scheduling method specifically designed for pose data characteristics. We also propose a masked training mechanism that effectively combines whole-body and part-specific datasets, enabling our model to capture interdependencies between body parts while avoiding overfitting to specific actions. Extensive experiments demonstrate DPoser-X's robustness and versatility across multiple benchmarks for body, hand, face, and full-body pose modeling. Our model consistently outperforms state-of-the-art alternatives, establishing a new benchmark for whole-body human pose prior modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。