用自监督学习构建通用触觉表征,让机器人更懂触摸。
Sparsh: Self-supervised touch representations for vision-based tactile sensing

- 通过掩码和自蒸馏在像素与潜在空间预训练,生成通用触觉表示。
- 在6个任务上平均性能比专用模型高95.1%,显著提升泛化能力。
- 适配多种触觉传感器,适合希望减少标注成本的研究者。
本文提出适用于日益普及的视觉触觉传感器的通用触觉表征方法。这类传感器虽显著补充了视觉感知,但现有方案多依赖任务与传感器特定的手工设计模型。由于真实数据采集需任务相关的标签(如接触力、滑动),且不同传感器在光照、凝胶标记等方面差异大,标注成本极高。为此,我们采用自监督学习(SSL),提出Sparsh系列模型,在460,000+张触觉图像上进行掩码与自蒸馏预训练。同时构建了TacBench基准测试平台,涵盖6项从触觉属性理解到物理感知与操作规划的任务。评估显示,基于SSL的预训练在TacBench上平均性能比端到端专用训练高出95.1%;其中Sparsh(DINO)与Sparsh(IJEPA)表现最优,表明潜在空间学习对触觉图像尤为有效。
原文摘要 · Abstract (English)
In this work, we introduce general purpose touch representations for the increasingly accessible class of vision-based tactile sensors. Such sensors have led to many recent advances in robot manipulation as they markedly complement vision, yet solutions today often rely on task and sensor specific handcrafted perception models. Collecting real data at scale with task centric ground truth labels, like contact forces and slip, is a challenge further compounded by sensors of various form factor differing in aspects like lighting and gel markings. To tackle this we turn to self-supervised learning (SSL) that has demonstrated remarkable performance in computer vision. We present Sparsh, a family of SSL models that can support various vision-based tactile sensors, alleviating the need for custom labels through pre-training on 460k+ tactile images with masking and self-distillation in pixel and latent spaces. We also build TacBench, to facilitate standardized benchmarking across sensors and models, comprising of six tasks ranging from comprehending tactile properties to enabling physical perception and manipulation planning. In evaluations, we find that SSL pre-training for touch representation outperforms task and sensor-specific end-to-end training by 95.1% on average over TacBench, and Sparsh (DINO) and Sparsh (IJEPA) are the most competitive, indicating the merits of learning in latent space for tactile images. Project page: https://sparsh-ssl.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。