arXiv:2606.30598cs.CV2026-06中稿 · ECCV

提出新数据集与模型,提升真实场景下手物3D姿态估计精度。

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

论文配图:Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
图 1 · 摘自论文原文
  • 用双向映射的接触数据构建新数据集EPIC-Contact
  • HOPformer模型在新数据集上成功率接近翻倍
  • 适合做具身智能与人机交互研究者参考

从真实场景下的第一视角RGB视频中准确估计3D手物姿态仍具挑战,主要源于严重遮挡和接触关系模糊。现有学习方法难以泛化到真实场景,且受限于标注稀缺。本文提出两项贡献:一是构建了包含2.3K个视频片段(共62.3K帧)的EPIC-Contact数据集,提供密集、一一对应的3D手物接触标注与姿态网格;二是提出HOPformer,一种端到端的Transformer模型,可单次前向传播联合预测双手与物体的姿态。其交叉注意力解码器以手部先验条件化物体特征,显著提升鲁棒性。在实验室数据集ARCTIC上,该模型达到82.4%的成功率,较当前最优提升6.2个百分点。在新提出的EPIC-Contact数据集上,成功率近乎翻倍,接触偏差降低75%。EPIC-Contact数据集、代码及模型权重已开源。

原文摘要 · Abstract (English)

Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are limited by the scarcity of supervision. We address these issues with two contributions. First, we introduce EPIC-Contact, an in-the-wild egocentric dataset of 2.3K clips (62.3K frames) with dense, bijective 3D hand-object contact correspondences and posed meshes. Second, we propose HOPformer, an end-to-end transformer that jointly predicts bi-manual hand and object pose in a single forward pass. A cross-attention decoder conditions object features on hand priors, producing robust pose estimation. We test HOPformer on the in-lab 3D dataset, ARCTIC, as well as our newly introduced EPIC-Contact dataset. HOPformer reaches 82.4% success rate on ARCTIC (+6.2 pts over current SOTA). On EPIC-Contact, it nearly doubles the success rate while reducing contact deviation by 75%. EPIC-Contact, HOPformer code and checkpoints are released: https://sid2697.github.io/epic-contact.

3D姿态估计第一视角手物交互数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。