arXiv:2510.00527cs.CV2025-10

用扩散模型分步生成手部3D姿态,同时捕捉不确定性。

Cascaded Diffusion Framework for Probabilistic Coarse-to-Fine Hand Pose Estimation

  • 先用扩散模型生成多种可能的关节位置,再用网格隐空间模型重建3D手形。
  • 在FreiHAND和HO3Dv2数据集上达到最新最好性能,能准确建模姿态分布。
  • 适合需要高精度且关心姿态不确定性的动作分析场景。

针对单阶段或级联式确定性方法在自遮挡与复杂手部结构下难以处理姿态歧义的问题,本文提出一种分阶段的扩散框架,结合概率建模与级联精修。第一阶段采用联合扩散模型采样多样化的3D关节假设;第二阶段利用网格隐空间扩散模型(Mesh LDM),基于采样关节构建3D手部网格。通过在学习的隐空间中使用多样化关节假设训练Mesh LDM,框架学会分布感知的关节-网格关系及鲁棒的手部先验。该级联设计缓解了直接从2D图像映射至密集3D姿态的困难,通过逐步精修提升准确性。在FreiHAND和HO3Dv2数据集上的实验表明,本方法不仅达到当前最优性能,还能有效建模姿态分布。

原文摘要 · Abstract (English)

Deterministic models for 3D hand pose reconstruction, whether single-staged or cascaded, struggle with pose ambiguities caused by self-occlusions and complex hand articulations. Existing cascaded approaches refine predictions in a coarse-to-fine manner but remain deterministic and cannot capture pose uncertainties. Recent probabilistic methods model pose distributions yet are restricted to single-stage estimation, which often fails to produce accurate 3D reconstructions without refinement. To address these limitations, we propose a coarse-to-fine cascaded diffusion framework that combines probabilistic modeling with cascaded refinement. The first stage is a joint diffusion model that samples diverse 3D joint hypotheses, and the second stage is a Mesh Latent Diffusion Model (Mesh LDM) that reconstructs a 3D hand mesh conditioned on a joint sample. By training Mesh LDM with diverse joint hypotheses in a learned latent space, our framework learns distribution-aware joint-mesh relationships and robust hand priors. Furthermore, the cascaded design mitigates the difficulty of directly mapping 2D images to dense 3D poses, enhancing accuracy through sequential refinement. Experiments on FreiHAND and HO3Dv2 demonstrate that our method achieves state-of-the-art performance while effectively modeling pose distributions.

3D手姿估计扩散模型概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。