arXiv:2604.00744cs.RO2026-04中稿 · publication at the…

用视觉变压器提升触觉感知,让机械手快速适配新传感器。

How to Train your Tactile Model: Tactile Perception with Multi-fingered Robot Hands

  • 采用视觉变压器提取触觉图像全局特征,避免依赖特定传感器数据。
  • 在五指机器人手上实现跨传感器泛化,接触属性识别准确率优于传统卷积网络。
  • 适合需要快速部署新触觉传感器的机器人场景,减少数据收集与重训练成本。

新触觉传感器的快速部署对可扩展的机器人操作至关重要,尤其在配备基于视觉的触觉传感器的多指机械手中。然而,当前推断接触特性的方法严重依赖卷积神经网络(CNN),虽然在已知传感器上表现良好,但需大量针对特定传感器的数据集,且因透镜特性、光照差异及传感器磨损等问题,每次更换传感器都需重新训练。本文提出TacViT,一种基于视觉变压器的新触觉感知模型,利用全局自注意力机制从触觉图像中提取鲁棒特征,实现对未见过传感器数据的准确接触属性推断。该能力显著减少了数据采集和重新训练的需求,加速了新传感器的部署。我们在五指机器人手上评估TacViT,结果表明其泛化性能优于现有CNN方法。研究凸显了TacViT在实际机器人应用中使触觉传感更具可扩展性和实用性方面的潜力。

原文摘要 · Abstract (English)

Rapid deployment of new tactile sensors is essential for scalable robotic manipulation, especially in multi-fingered hands equipped with vision-based tactile sensors. However, current methods for inferring contact properties rely heavily on convolutional neural networks (CNNs), which, while effective on known sensors, require large, sensor-specific datasets. Furthermore, they require retraining for each new sensor due to differences in lens properties, illumination, and sensor wear. Here we introduce TacViT, a novel tactile perception model based on Vision Transformers, designed to generalize on new sensor data. TacViT leverages global self-attention mechanisms to extract robust features from tactile images, enabling accurate contact property inference even on previously unseen sensors. This capability significantly reduces the need for data collection and retraining, accelerating the deployment of new sensors. We evaluate TacViT on sensors for a five-fingered robot hand and demonstrate its superior generalization performance compared to CNNs. Our results highlight TacViTs potential to make tactile sensing more scalable and practical for real-world robotic applications.

触觉感知视觉变压器机器人泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。