arXiv:2506.14754cs.RO2025-06被引 21

多模态触觉表征让机器人更懂触摸,提升抓取成功率63%

Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation

  • 融合图像、声音、运动、压力四类触觉信号,构建统一表征
  • 在100万次交互数据上预训练,使抓取成功率提升63%
  • 适合需要精细触觉反馈的机器人操作任务研究者

我们提出Sparsh-X,首个整合四种触觉模态(图像、音频、运动、压力)的多模态触觉表征。基于在Digit 360传感器上收集的约100万次丰富接触交互数据训练,Sparsh-X在不同时空尺度下捕捉互补的触觉信号。通过自监督学习,将多模态信息融合为统一表征,有效捕获对机器人操作有用物理属性。研究了如何将真实触觉表征用于模仿学习和仿真策略的触觉适应,结果显示,相比仅使用触觉图像的端到端模型,Sparsh-X使策略成功率提升63%,在恢复物体状态时鲁棒性提高90%。此外,我们在物体-动作识别、材料-数量估计、力值估计等任务上进行基准测试,发现相较于端到端方法,Sparsh-X在表征物理属性方面准确率提升48%,验证了多模态预训练在捕捉灵巧操作关键特征上的优势。

原文摘要 · Abstract (English)

We present Sparsh-X, the first multisensory touch representations across four tactile modalities: image, audio, motion, and pressure. Trained on ~1M contact-rich interactions collected with the Digit 360 sensor, Sparsh-X captures complementary touch signals at diverse temporal and spatial scales. By leveraging self-supervised learning, Sparsh-X fuses these modalities into a unified representation that captures physical properties useful for robot manipulation tasks. We study how to effectively integrate real-world touch representations for both imitation learning and tactile adaptation of sim-trained policies, showing that Sparsh-X boosts policy success rates by 63% over an end-to-end model using tactile images and improves robustness by 90% in recovering object states from touch. Finally, we benchmark Sparsh-X ability to make inferences about physical properties, such as object-action identification, material-quantity estimation, and force estimation. Sparsh-X improves accuracy in characterizing physical properties by 48% compared to end-to-end approaches, demonstrating the advantages of multisensory pretraining for capturing features essential for dexterous manipulation.

触觉表征多模态机器人操作自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。