arXiv:2506.15953cs.RO2025-06被引 53

融合视觉与触觉信息,提升机器人精细操作成功率。

ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

  • 用交叉注意力融合高分辨率视觉与触觉数据。
  • 在真实场景中实现50%更高的操作成功率。
  • 首次让机械手自主完成11步连续精细操作。

精细操作是机器人系统实现类人物理交互的核心能力。尽管基于视觉的方法发展迅速,但在非结构化或视觉遮挡环境下,触觉感知对细微控制仍至关重要。我们提出ViTacFormer,一种将交叉注意力编码器与自回归触觉预测头结合的表征学习方法,用于融合高分辨率视觉与触觉信号。在此基础上,设计由易到难的渐进式训练策略,逐步优化视觉-触觉潜在空间,提升模型精度与鲁棒性。所学跨模态表征支持多指机械手的模仿学习,实现精准自适应操作。在一系列复杂真实世界基准测试中,本方法相比先前最先进系统成功率提升约50%。据我们所知,这也是首个能自主完成需高精度控制的长时程精细操作任务的系统,成功执行最多11个连续阶段,并持续运行2.5分钟。

原文摘要 · Abstract (English)

Dexterous manipulation is a cornerstone capability for robotic systems aiming to interact with the physical world in a human-like manner. Although vision-based methods have advanced rapidly, tactile sensing remains crucial for fine-grained control, particularly in unstructured or visually occluded settings. We present ViTacFormer, a representation-learning approach that couples a cross-attention encoder to fuse high-resolution vision and touch with an autoregressive tactile prediction head that anticipates future contact signals. Building on this architecture, we devise an easy-to-challenging curriculum that steadily refines the visual-tactile latent space, boosting both accuracy and robustness. The learned cross-modal representation drives imitation learning for multi-fingered hands, enabling precise and adaptive manipulation. Across a suite of challenging real-world benchmarks, our method achieves approximately 50% higher success rates than prior state-of-the-art systems. To our knowledge, it is also the first to autonomously complete long-horizon dexterous manipulation tasks that demand highly precise control with an anthropomorphic hand, successfully executing up to 11 sequential stages and sustaining continuous operation for 2.5 minutes.

机器人操作跨模态学习触觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。