提出统一高效框架,用扩散模型分离建模骨骼结构,提升3D人体姿态估计精度与泛化能力。
FastDDHPose: Towards Unified, Efficient, and Disentangled 3D Human Pose Estimation
- 基于扩散模型分离建模骨长与骨向分布,避免层级误差累积。
- 设计轻量级运动-层级时空去噪器,聚焦关节层次关系,减少复杂拓扑建模。
- 在Human3.6M和MPI-INF-3DHP上实现领先性能,训练效率显著提升。
单目3D人体姿态估计近年通过直接从2D关键点序列回归3D姿态取得显著进展。然而,现有方法多在不同框架下训练与评估,缺乏统一比较标准。为此,我们提出Fast3DHPE,一个模块化框架,支持快速复现与灵活开发新方法。通过标准化训练与评估流程,该框架实现方法间公平比较,并大幅提高训练效率。在此框架内,我们提出FastDDHPose,一种基于扩散模型的解耦3D人体姿态估计方法,利用扩散模型强大的潜在分布建模能力,显式建模骨长与骨向分布,避免层级误差放大。同时,设计高效的运动-层级时空去噪器,引导模型关注关节层次结构,减少对复杂拓扑的冗余建模。在Human3.6M和MPI-INF-3DHP上的大量实验表明,该框架实现方法间公平比较,显著提升训练效率;在此统一框架下,FastDDHPose在真实场景中展现卓越泛化性与鲁棒性,达到当前最优性能。代码与模型将开源:https://github.com/Andyen512/Fast3DHPE。
原文摘要 · Abstract (English)
Recent approaches for monocular 3D human pose estimation (3D HPE) have achieved leading performance by directly regressing 3D poses from 2D keypoint sequences. Despite the rapid progress in 3D HPE, existing methods are typically trained and evaluated under disparate frameworks, lacking a unified framework for fair comparison. To address these limitations, we propose Fast3DHPE, a modular framework that facilitates rapid reproduction and flexible development of new methods. By standardizing training and evaluation protocols, Fast3DHPE enables fair comparison across 3D human pose estimation methods while significantly improving training efficiency. Within this framework, we introduce FastDDHPose, a Disentangled Diffusion-based 3D Human Pose Estimation method which leverages the strong latent distribution modeling capability of diffusion models to explicitly model the distributions of bone length and bone direction while avoiding further amplification of hierarchical error accumulation. Moreover, we design an efficient Kinematic-Hierarchical Spatial and Temporal Denoiser that encourages the model to focus on kinematic joint hierarchies while avoiding unnecessary modeling of overly complex joint topologies. Extensive experiments on Human3.6M and MPI-INF-3DHP show that the Fast3DHPE framework enables fair comparison of all methods while significantly improving training efficiency. Within this unified framework, FastDDHPose achieves state-of-the-art performance with strong generalization and robustness in in-the-wild scenarios. The framework and models will be released at: https://github.com/Andyen512/Fast3DHPE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。