从视频推断双手触觉,让机器学会真实交互的物理感受。
TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video

- 用多视角视频+穿戴触觉传感器,构建大规模双手交互数据集。
- 结合手腕摄像头能提升触觉预测准确率,最高提升6.1%。
- 适合研究具身智能、人机交互与视觉触觉融合的学者。
以第一视角捕捉人类-环境交互的视频数据已成为推动具身智能研究的关键,但现有数据集普遍缺乏触觉信号,而高精度触觉硬件部署成本高昂。本文提出EgoTouch,一个大规模多视角第一人称数据集,涵盖1,891个场景中的208种操作任务,包含头戴式与双腕摄像机的同步RGB视频、双手3D姿态及可穿戴触觉传感器生成的连续压力图。基于此,我们构建了TouchAnything框架,以第一视角为主输入,灵活利用腕部摄像头信息进行触觉预测。实验表明,引入腕部视角可显著提升预测性能,接触交并比(Contact IoU)提升5.0%,体积交并比(Volumetric IoU)提升6.1%。数据集、代码与基准将公开发布。
原文摘要 · Abstract (English)
Egocentric human video data, which captures rich human-environment interactions and can be collected at scale, has become a key driver of embodied intelligence research. However, existing egocentric datasets typically lack tactile sensing, a critical modality that provides direct cues about contact, force, and pressure in human-object interaction. Without such signals, models struggle to learn physically grounded representations of real-world interaction dynamics. While tactile sensors provide these cues, deploying high-quality tactile hardware at scale remains expensive and cumbersome. This raises a central question: can tactile feedback be inferred directly from visual observations, enabling scalable tactile supervision for egocentric video data and supporting physically grounded embodied learning? To enable research in this direction, we introduce EgoTouch, a large-scale multi-view egocentric dataset with dense tactile supervision for bimanual hand-object interaction. EgoTouch comprises 208 manipulation tasks spanning 1,891 episodes in diverse indoor and outdoor environments, with synchronized multi-view RGB (head-mounted egocentric and dual wrist-mounted cameras), bimanual 3D hand pose, and continuous pressure maps from wearable tactile sensors. Building on EgoTouch, we introduce TouchAnything, a baseline multi-view vision-to-touch prediction framework that uses the egocentric view as the primary input and flexibly leverages available wrist-mounted views at inference time. Experiments show that incorporating wrist-mounted views generally improves tactile prediction over egocentric-only input, achieving up to 5.0% relative improvement in Contact IoU and 6.1% relative improvement in Volumetric IoU. We will publicly release the dataset, code, and benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。