arXiv:2606.06281cs.RO2026-06

融合多分辨率触觉传感器,提升机器人复杂操作任务成功率。

Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

论文配图:Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation
图 1 · 摘自论文原文
  • 设计多传感器融合架构,整合视觉、凝胶触觉与事件触觉数据。
  • 在5个高接触任务中平均成功率80%,远超纯视觉(31%)和视觉触觉(54%)基线。
  • 无需事件触觉传感器即可提升性能,适合实际部署的机器人学习应用。

触觉感知对解决多种操作任务有益。尽管存在多种具有不同特性的触觉传感器,但利用异构触觉传感器融合来提升操作学习仍研究不足。本文提出多分辨率触觉感知(MiTaS)框架,通过不同时间分辨率的多个触觉传感器协同解决复杂的高接触操作任务。我们设计了一种新型架构,采用模态特定卷积主干与基于Transformer的融合机制,有效融合来自RGB相机流、基于视觉的GelSight Mini传感器和高频事件型Evetac传感器的数据。该多传感器表征用于条件化流匹配策略以完成下游任务。在五个高接触操作任务上的实验结果表明,多分辨率触觉特征在模仿学习中具有显著有效性:MiTaS平均成功率达80%,而仅使用视觉(31%)和视觉-触觉(54%)的基线无法可靠完成任务。使用多触觉数据联合训练的视觉触觉模型,在某些任务上性能提升超过10%,即使在策略评估时未接入Evetac传感器也有效。传感器读数与注意力分析揭示了各传感器在任务执行过程中的关键作用,验证了多分辨率触觉感知的有效性。

原文摘要 · Abstract (English)

Touch sensing is beneficial for solving a wide variety of manipulation tasks. While there exists a wide range of tactile sensors with different properties, exploiting the fusion of multiple heterogeneous tactile sensors to improve manipulation learning remains underexplored. We present Multi-Resolution Tactile Sensing (MiTaS), a representation framework that leverages multiple tactile sensors operating at different temporal resolutions in order to solve complex contact-rich manipulation tasks. We propose a novel architecture using modality-specific convolutional stems and transformer-based fusion that effectively fuses information from an RGB camera stream, a vision-based GelSight Mini sensor and a high-frequency event-based Evetac sensor. This multi-sensor representation then conditions a flow-matching policy for solving downstream tasks. Experimental results across five contact-rich manipulation tasks demonstrate the effectiveness of multi-resolution tactile features in imitation learning. MiTaS achieves an average success rate of 80 %, while vision-only (31 %) and visual-tactile (54 %) baselines cannot solve the task reliably. Co-training a visuo-tactile model with multi-tactile data boosts performance by over 10 \% in certain tasks, without having access to the Evetac sensor during policy evaluation. A detailed sensor-reading and attention analysis reveals the importance of different sensors throughout task execution, validating our multi-resolution tactile sensing approach. Project Page: http://mitas-touch.github.io.

触觉感知机器人操作多传感器融合模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。