用超图扩散模型提升单目3D人体姿态估计精度。
HyperDiff: Hypergraph Guided Diffusion Model for 3D Human Pose Estimation
- 结合扩散模型与超图网络,建模关节间高阶关系。
- 在Human3.6M和MPI-INF-3DHP上达当前最佳性能。
- 可灵活适配不同算力,兼顾效果与效率。
单目3D人体姿态估计在从2D到3D映射过程中常面临深度模糊和遮挡问题。传统方法在利用骨骼结构信息时可能忽略多尺度特征,影响估计精度。本文提出HyperDiff,将扩散模型与超图卷积网络(HyperGCN)结合:扩散模型有效捕捉数据不确定性,缓解深度模糊与遮挡;超图网络作为去噪器,通过多粒度结构精准建模关节间的高阶相关性,显著提升复杂姿态下的去噪能力。实验表明,HyperDiff在Human3.6M和MPI-INF-3DHP数据集上均达到当前最优性能,并能灵活适应不同计算资源,在性能与效率间实现良好平衡。
原文摘要 · Abstract (English)
Monocular 3D human pose estimation (HPE) often encounters challenges such as depth ambiguity and occlusion during the 2D-to-3D lifting process. Additionally, traditional methods may overlook multi-scale skeleton features when utilizing skeleton structure information, which can negatively impact the accuracy of pose estimation. To address these challenges, this paper introduces a novel 3D pose estimation method, HyperDiff, which integrates diffusion models with HyperGCN. The diffusion model effectively captures data uncertainty, alleviating depth ambiguity and occlusion. Meanwhile, HyperGCN, serving as a denoiser, employs multi-granularity structures to accurately model high-order correlations between joints. This improves the model's denoising capability especially for complex poses. Experimental results demonstrate that HyperDiff achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets and can flexibly adapt to varying computational resources to balance performance and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。