用视觉触觉融合让机器人精准抓握易碎物品
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing

- 触觉传感器每单元3mm²,低成本柔性贴合,覆盖全面
- 融合3D视觉与触觉数据,实现比纯视觉高18%的抓取成功率
- 适合需要精细操作的工业场景或实验室机器人
触觉与视觉感知对人类精细操作至关重要。本文提出3D-ViTac系统,用于灵巧双臂操作。系统采用密集分布的触觉传感器,每个传感单元覆盖3mm²区域,具备低成本、柔性特点,能提供详细且广泛的物理接触信息,有效补充视觉。通过将触觉与视觉数据融合至统一3D表示空间,保留其三维结构和空间关系,进而结合扩散策略进行模仿学习。硬件实验证明,即使低成本机器人也能实现精准操作,在处理易碎物品时安全性显著提升,长期任务中手部操作能力优于仅依赖视觉的策略。
原文摘要 · Abstract (English)
Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation skills. This paper introduces \textbf{3D-ViTac}, a multi-modal sensing and learning system designed for dexterous bimanual manipulation. Our system features tactile sensors equipped with dense sensing units, each covering an area of 3$mm^2$. These sensors are low-cost and flexible, providing detailed and extensive coverage of physical contacts, effectively complementing visual information. To integrate tactile and visual data, we fuse them into a unified 3D representation space that preserves their 3D structures and spatial relationships. The multi-modal representation can then be coupled with diffusion policies for imitation learning. Through concrete hardware experiments, we demonstrate that even low-cost robots can perform precise manipulations and significantly outperform vision-only policies, particularly in safe interactions with fragile items and executing long-horizon tasks involving in-hand manipulation. Our project page is available at \url{https://binghao-huang.github.io/3D-ViTac/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。