多模态触觉表征让机器人更懂触摸,提升抓取成功率63%
Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation
- 融合图像、声音、运动、压力四类触觉信号,构建统一表征
- 在100万次交互数据上预训练,使抓取成功率提升63%
- 适合需要精细触觉反馈的机器人操作任务研究者
我们提出Sparsh-X,首个整合四种触觉模态(图像、音频、运动、压力)的多模态触觉表征。基于在Digit 360传感器上收集的约100万次丰富接触交互数据训练,Sparsh-X在不同时空尺度下捕捉互补的触觉信号。通过自监督学习,将多模态信息融合为统一表征,有效捕获对机器人操作有用物理属性。研究了如何将真实触觉表征用于模仿学习和仿真策略的触觉适应,结果显示,相比仅使用触觉图像的端到端模型,Sparsh-X使策略成功率提升63%,在恢复物体状态时鲁棒性提高90%。此外,我们在物体-动作识别、材料-数量估计、力值估计等任务上进行基准测试,发现相较于端到端方法,Sparsh-X在表征物理属性方面准确率提升48%,验证了多模态预训练在捕捉灵巧操作关键特征上的优势。
原文摘要 · Abstract (English)
We present Sparsh-X, the first multisensory touch representations across four tactile modalities: image, audio, motion, and pressure. Trained on ~1M contact-rich interactions collected with the Digit 360 sensor, Sparsh-X captures complementary touch signals at diverse temporal and spatial scales. By leveraging self-supervised learning, Sparsh-X fuses these modalities into a unified representation that captures physical properties useful for robot manipulation tasks. We study how to effectively integrate real-world touch representations for both imitation learning and tactile adaptation of sim-trained policies, showing that Sparsh-X boosts policy success rates by 63% over an end-to-end model using tactile images and improves robustness by 90% in recovering object states from touch. Finally, we benchmark Sparsh-X ability to make inferences about physical properties, such as object-action identification, material-quantity estimation, and force estimation. Sparsh-X improves accuracy in characterizing physical properties by 48% compared to end-to-end approaches, demonstrating the advantages of multisensory pretraining for capturing features essential for dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。