利用手背皮肤形变提升第一人称视角下手指遮挡时的姿态估计精度
DeltaDorsal: Enhancing Hand Pose Estimation with Dorsal Features in Egocentric Views
- 通过对比动态手势与放松姿态的特征,构建双流增量编码器
- 在手指遮挡超50%时,平均关节误差降低18%
- 适合轻量级设备部署,支持无视觉运动的力反馈交互
XR设备的普及使第一人称视角的手部姿态估计成为关键任务,但该视角常因手指频繁遮挡而面临挑战。为此,我们提出一种新方法,利用近期密集视觉特征提取器解锁的手背皮肤形变丰富信息。引入双流增量编码器,通过对比动态手部与基准放松状态的特征来学习姿态。评估显示,仅使用裁剪后的手背图像,在手指遮挡比例≥50%的情况下,我们的方法将平均每关节角度误差(MPJAE)相比依赖完整手部几何和大模型骨干的最先进方法降低了18%。因此,该方法不仅提升了遮挡场景下指尖捏合与点击等下游任务的可靠性,还开启了新的交互范式,例如在无可见运动时检测等张力以实现表面“点击”,同时显著减小模型规模。
原文摘要 · Abstract (English)
The proliferation of XR devices has made egocentric hand pose estimation a vital task, yet this perspective is inherently challenged by frequent finger occlusions. To address this, we propose a novel approach that leverages the rich information in dorsal hand skin deformation, unlocked by recent advances in dense visual featurizers. We introduce a dual-stream delta encoder that learns pose by contrasting features from a dynamic hand with a baseline relaxed position. Our evaluation demonstrates that, using only cropped dorsal images, our method reduces the Mean Per Joint Angle Error (MPJAE) by 18% in self-occluded scenarios (fingers >= 50% occluded) compared to state-of-the-art techniques that depend on the whole hand's geometry and large model backbones. Consequently, our method not only enhances the reliability of downstream tasks like index finger pinch and tap estimation in occluded scenarios but also unlocks new interaction paradigms, such as detecting isometric force for a surface "click" without visible movement while minimizing model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。