arXiv:2605.17742cs.CVcs.HC2026-05中稿 · CVPR被引 1

通过不确定性建模提升自监督手部姿态估计的稳定性与精度

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

论文配图:UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
图 1 · 摘自论文原文
  • 引入概率点云空间,用条件归一化流捕捉姿态分布
  • 在三个数据集上将误差降低37.8%,显著优于现有方法
  • 适合对自监督三维姿态估计和鲁棒性训练有需求的研究者

人工标注精确的3D手部姿态极其耗时。现有自监督方法依赖输入图像与渲染输出的差异或多视图一致性作为优化驱动力,但易受噪声伪标签影响,且忽视细粒度空间相关性,导致训练不稳定。为此,我们提出UST-Hand,一种自监督学习框架,通过估计手部姿态的不确定性分布并构建概率点云特征空间,实现复杂时空关系建模。该框架采用条件归一化流模型捕捉姿态分布,生成多样化假设,增强在噪声伪标签下的鲁棒性。这些多假设被映射至统一的概率3D点云空间,实现多视图与时间特征交互,全面探索手部运动模式与细粒度空间相关性。在三个挑战性数据集上的实验表明,UST-Hand达到领先性能,其平均顶点位置误差(MPVPE)相比现有方法最高提升37.8%。

原文摘要 · Abstract (English)

Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and progressively refine pose accuracy. However, these methods are highly susceptible to noisy pseudo-labels and overlook the importance of fully exploiting fine-grained spatial correlations, which undermines the stability of model training. To address these issues, we propose UST-Hand, a self-supervised learning framework that estimates uncertainty distribution of hand pose and constructs a probabilistic point cloud feature space, which enables the complex spatiotemporal relationship modeling. UST-Hand employs a conditional normalizing flow model to capture hand pose distributions and samples diverse hypotheses, facilitating robust learning under noisy pseudo-labels supervision with enhanced stability. These multi-hypothesis are mapped to a unified probabilistic 3D point cloud space for multi-view and temporal feature interaction, comprehensively exploring hand motion patterns and fine-grained spatial correlations. Extensive experiments on three challenging datasets demonstrate that UST-Hand achieves state-of-the-art performance, outperforming existing self-supervised methods by up to 37.8% in Mean Per Vertex Position Error (MPVPE).

自监督学习3D姿态估计不确定性建模点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。