arXiv:2508.14431cs.CV2025-08被引 2

用超图扩散模型提升单目3D人体姿态估计精度。

HyperDiff: Hypergraph Guided Diffusion Model for 3D Human Pose Estimation

  • 结合扩散模型与超图网络,建模关节间高阶关系。
  • 在Human3.6M和MPI-INF-3DHP上达当前最佳性能。
  • 可灵活适配不同算力,兼顾效果与效率。

单目3D人体姿态估计在从2D到3D映射过程中常面临深度模糊和遮挡问题。传统方法在利用骨骼结构信息时可能忽略多尺度特征,影响估计精度。本文提出HyperDiff,将扩散模型与超图卷积网络(HyperGCN)结合:扩散模型有效捕捉数据不确定性,缓解深度模糊与遮挡;超图网络作为去噪器,通过多粒度结构精准建模关节间的高阶相关性,显著提升复杂姿态下的去噪能力。实验表明,HyperDiff在Human3.6M和MPI-INF-3DHP数据集上均达到当前最优性能,并能灵活适应不同计算资源,在性能与效率间实现良好平衡。

原文摘要 · Abstract (English)

Monocular 3D human pose estimation (HPE) often encounters challenges such as depth ambiguity and occlusion during the 2D-to-3D lifting process. Additionally, traditional methods may overlook multi-scale skeleton features when utilizing skeleton structure information, which can negatively impact the accuracy of pose estimation. To address these challenges, this paper introduces a novel 3D pose estimation method, HyperDiff, which integrates diffusion models with HyperGCN. The diffusion model effectively captures data uncertainty, alleviating depth ambiguity and occlusion. Meanwhile, HyperGCN, serving as a denoiser, employs multi-granularity structures to accurately model high-order correlations between joints. This improves the model's denoising capability especially for complex poses. Experimental results demonstrate that HyperDiff achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets and can flexibly adapt to varying computational resources to balance performance and efficiency.

3D姿态估计扩散模型超图网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。