首个真实场景全手触觉数据集,让机器理解人手如何触摸物体。
OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- 构建首个野外真实场景的全手触觉视频数据集
- 触觉信号能有效提升抓握理解与跨模态对齐效果
- 适合研究具身智能、多模态感知与机器人操作的人看
人类手是与物理世界交互的主要工具,但第一人称视觉通常无法感知接触的时间、位置和力度。现有的可穿戴触觉传感器性能有限,且缺乏同步的第一人称视频与全手触觉数据集。为弥合视觉感知与物理交互之间的差距,我们提出了 OpenTouch,首个在真实场景中采集的全手触觉数据集,包含5.1小时同步的视频-触觉-姿态数据,以及2,900段带详细文本标注的精选片段。基于 OpenTouch,我们建立了检索与分类基准,探究触觉如何支撑感知与行为。实验表明,触觉信号是抓握理解的紧凑而有力线索,能增强跨模态对齐,并可从真实场景视频中可靠检索。通过公开此标注的视觉-触觉-姿态数据集与基准,我们旨在推动多模态第一人称感知、具身学习与高接触密度的机器人操作研究。
原文摘要 · Abstract (English)
The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wild datasets align first-person video with full-hand touch. To bridge the gap between visual perception and physical interaction, we present OpenTouch, the first in-the-wild egocentric full-hand tactile dataset, containing 5.1 hours of synchronized video-touch-pose data and 2,900 curated clips with detailed text annotations. Using OpenTouch, we introduce retrieval and classification benchmarks that probe how touch grounds perception and action. We show that tactile signals provide a compact yet powerful cue for grasp understanding, strengthen cross-modal alignment, and can be reliably retrieved from in-the-wild video queries. By releasing this annotated vision-touch-pose dataset and benchmark, we aim to advance multimodal egocentric perception, embodied learning, and contact-rich robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。