构建多模态操作数据集,融合视觉与力觉感知,支持人机交互研究。
Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- 采集381个可动物体在38环境中的3048组交互序列,含四种操作形态。
- 首次将力反馈与多视角视频同步,覆盖人类手、带摄像头的手腕、机械臂等形态。
- 适合研究机器人抓取、人机协作及跨视角迁移的学者使用。
我们提出一个面向力觉感知、跨视角可动物体操作的数据集,将真实人机交互中的视觉、动作与触觉感受进行耦合。数据集包含381个可动物体在38种环境中的3048个交互序列。每个物体在四种操作形态下被操作:(i) 人类手,(ii) 带腕部相机的人类手,(iii) 手持UMI夹爪,(iv) 自定义Hoi!夹爪,其中工具形态提供末端执行器的力信号与触觉传感。该数据集从视频角度提供交互理解的全景视图,使研究者可评估方法在人与机器人视角间的迁移性能,并探索交互力等未充分研究的模态。项目主页见 https://timengelbracht.github.io/Hoi-Dataset-Website/。
原文摘要 · Abstract (English)
We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments - (i) human hand, (ii) human hand with a wrist-mounted camera, (iii) handheld UMI gripper, and (iv) a custom Hoi! gripper, where the tool embodiment provides end-effector forces and tactile sensing. Our dataset offers a holistic view of interaction understanding from video, enabling researchers to evaluate how well methods transfer between human and robotic viewpoints, but also investigate underexplored modalities such as interaction forces. The Project Website can be found at https://timengelbracht.github.io/Hoi-Dataset-Website/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。