融合方向感知特征,提升3D手部姿态估计精度与稳定性
Direction-Aware Hybrid Representation Learning for 3D Hand Pose and Shape Estimation
- 用隐式图像特征和显式2D坐标特征融合,引入像素方向信息
- 在FreiHAND上准确率比现有方法高33%,HO3Dv2/v3榜单领先10%
- 适合实时动捕场景,抗遮挡、模糊和手位变化
大多数基于模型的3D手部姿态与形状估计方法直接从图像回归参数化模型参数,在弱监督下获得3D关节位置。然而,这类方法需解决具有多个局部极小值的复杂优化问题,训练困难。为此,我们提出方向感知混合特征(DaHyF),融合隐式图像特征与显式2D关节坐标特征,并利用相机坐标系中的像素方向信息来估计姿态、形状和相机视角。该方法通过DaHyF表示直接预测3D手部姿态,并基于对比学习预测置信度,减少运动捕捉中的抖动。我们在FreiHAND数据集上评估,准确率超越现有最先进方法超过33%。在HO3Dv2和HO3Dv3排行榜上,于均关节误差(经尺度与平移对齐后)指标上排名第一,相比第二名最大提升达10%。此外,该方法在手位变化、遮挡和运动模糊的真实场景中也表现出色。
原文摘要 · Abstract (English)
Most model-based 3D hand pose and shape estimation methods directly regress the parametric model parameters from an image to obtain 3D joints under weak supervision. However, these methods involve solving a complex optimization problem with many local minima, making training difficult. To address this challenge, we propose learning direction-aware hybrid features (DaHyF) that fuse implicit image features and explicit 2D joint coordinate features. This fusion is enhanced by the pixel direction information in the camera coordinate system to estimate pose, shape, and camera viewpoint. Our method directly predicts 3D hand poses with DaHyF representation and reduces jittering during motion capture using prediction confidence based on contrastive learning. We evaluate our method on the FreiHAND dataset and show that it outperforms existing state-of-the-art methods by more than 33% in accuracy. DaHyF also achieves the top ranking on both the HO3Dv2 and HO3Dv3 leaderboards for the metric of Mean Joint Error (after scale and translation alignment). Compared to the second-best results, the largest improvement observed is 10%. We also demonstrate its effectiveness in real-time motion capture scenarios with hand position variability, occlusion, and motion blur.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。