用便携触觉夹爪收集真实场景的多模态数据,提升机器人精细操作能力。
Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile Gripper
- 设计可穿戴触觉夹爪,同步采集视觉与触觉数据。
- 跨模态学习让模型聚焦接触区域,提升操作精度。
- 在试管插入等任务中表现更稳健,抗干扰能力强。
手持式夹爪因部署便捷、适用性强,被广泛用于收集人类操作示范。然而,现有设计普遍缺乏触觉传感,而触觉反馈对精准操作至关重要。本文提出一种轻量化、便携式夹爪,集成触觉传感器,可在多样真实环境中同步采集视觉与触觉数据。基于此硬件,我们构建了一种跨模态表示学习框架,融合视觉与触觉信号,同时保留各自特性。该学习过程促使模型生成可解释的表示,始终聚焦于物理交互相关的接触区域。在下游操作任务中,这些表示显著提升策略学习效率与效果,支持基于多模态反馈的精准机器人操作。我们在试管插入和移液器液体转移等细粒度任务上验证方法,结果表明在外部扰动下仍具备更高准确率与鲁棒性。
原文摘要 · Abstract (English)
Handheld grippers are increasingly used to collect human demonstrations due to their ease of deployment and versatility. However, most existing designs lack tactile sensing, despite the critical role of tactile feedback in precise manipulation. We present a portable, lightweight gripper with integrated tactile sensors that enables synchronized collection of visual and tactile data in diverse, real-world, and in-the-wild settings. Building on this hardware, we propose a cross-modal representation learning framework that integrates visual and tactile signals while preserving their distinct characteristics. The learning procedure allows the emergence of interpretable representations that consistently focus on contacting regions relevant for physical interactions. When used for downstream manipulation tasks, these representations enable more efficient and effective policy learning, supporting precise robotic manipulation based on multimodal feedback. We validate our approach on fine-grained tasks such as test tube insertion and pipette-based fluid transfer, demonstrating improved accuracy and robustness under external disturbances. Our project page is available at https://binghao-huang.github.io/touch_in_the_wild/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。