arXiv:2603.28760cs.CVcs.RO2026-03被引 3

首个在真实场景中捕捉手物交互3D数据的开源数据集

SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild

  • 无标记多相机系统,配合VR头显实现自由移动采集
  • 构建首个包含户外等真实环境的3D手物交互数据集
  • 适合研究真实场景下人机交互与3D感知的科研人员

当前基于第一视角的计算机视觉在理解人类手部与物体交互时仍面临挑战。现有手物交互数据集多在受控摄影棚内采集,限制了环境多样性,也影响模型在真实场景中的泛化能力。为此,我们提出一种无标记的多相机系统,可实现几乎不受限的野外移动采集,同时保持高精度3D标注能力。该系统由轻量化背挂式多相机阵列构成,与用户佩戴的VR头显同步校准。针对3D真实标注,我们开发了自外向内跟踪流水线,并严格评估其质量。最终发布SHOW3D,这是首个大规模、具有3D标注的真实世界手物交互数据集,涵盖多种真实环境(包括户外)。实验验证了该方法显著缓解了环境真实性与3D标注精度之间的根本矛盾。

原文摘要 · Abstract (English)

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings, which limits both environmental diversity and the ability of models trained on such data to generalize to real-world scenarios. To address this challenge, we introduce a novel marker-less multi-camera system that allows for nearly unconstrained mobility in genuinely in-the-wild conditions, while still having the ability to generate precise 3D annotations of hands and objects. The capture system consists of a lightweight, back-mounted, multi-camera rig that is synchronized and calibrated with a user-worn VR headset. For 3D ground-truth annotation of hands and objects, we develop an ego-exo tracking pipeline and rigorously evaluate its quality. Finally, we present SHOW3D, the first large-scale dataset with 3D annotations that show hands interacting with objects in diverse real-world environments, including outdoor settings. Our approach significantly reduces the fundamental trade-off between environmental realism and accuracy of 3D annotations, which we validate with experiments on several downstream tasks. show3d-dataset.github.io

3D感知手物交互数据集真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。