arXiv:2603.17396cs.CV2026-03

利用手势语义提升单目图像3D手部姿态估计精度

Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation

  • 用粗细手势标签进行预训练,构建语义感知的嵌入空间
  • 在InterHand2.6M上比EANet基线提升单手姿态准确率
  • 方法可跨模型迁移,无需修改即可提升性能

从单目RGB图像估计3D手部姿态是AR/VR、人机交互和手语理解中的关键任务。本文聚焦于存在离散手势标签的场景,证明手势语义可作为3D姿态估计的强大先验。提出两阶段框架:首先在InterHand2.6M数据集上使用粗粒度和细粒度手势标签进行手势感知预训练,学习信息丰富的嵌入空间;随后通过基于手势嵌入的逐关节令牌Transformer,回归MANO手部参数。训练采用分层目标函数,优化参数、关节及结构约束。在InterHand2.6M上的实验表明,该方法持续优于当前最优的EANet基线,且效果可跨不同架构迁移,无需修改。

原文摘要 · Abstract (English)

Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available and show that gesture semantics can serve as a powerful inductive bias for 3D pose estimation. We present a two-stage framework: gesture-aware pretraining that learns an informative embedding space using coarse and fine gesture labels from InterHand2.6M, followed by a per-joint token Transformer guided by gesture embeddings as intermediate representations for final regression of MANO hand parameters. Training is driven by a layered objective over parameters, joints, and structural constraints. Experiments on InterHand2.6M demonstrate that gesture-aware pretraining consistently improves single-hand accuracy over the state-of-the-art EANet baseline, and that the benefit transfers across architectures without any modification.

3D手部姿态手势语义预训练Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。