融合视觉触觉信息,提升机器人在遮挡场景下的操作能力
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
- 用软融合对比学习增强多模态特征对齐与互补性
- 在12个仿真和3个真实环境上显著优于基线方法
- 适合需要多传感器融合的机器人操控研究者
触觉信息在人类操作任务中至关重要,近年来在机器人操作领域受到越来越多关注。然而,现有方法大多聚焦于视觉与触觉特征的对齐,融合机制多为直接拼接,忽视了两种模态间的内在互补性,且对齐利用不足,导致在遮挡场景下表现不佳,限制了实际应用潜力。本文提出ViTaS,一种简单而有效的框架,通过引入软融合对比学习(Soft Fusion Contrastive Learning)和条件变分自编码器(CVAE)模块,充分利用视觉-触觉表示中的对齐关系与互补特性,指导智能体行为。我们在12个模拟环境和3个真实世界环境中验证了该方法的有效性,实验结果表明,ViTaS显著优于现有基线方法。
原文摘要 · Abstract (English)
Tactile information plays a crucial role in human manipulation tasks and has recently garnered increasing attention in robotic manipulation. However, existing approaches mostly focus on the alignment of visual and tactile features and the integration mechanism tends to be direct concatenation. Consequently, they struggle to effectively cope with occluded scenarios due to neglecting the inherent complementary nature of both modalities and the alignment may not be exploited enough, limiting the potential of their real-world deployment. In this paper, we present ViTaS, a simple yet effective framework that incorporates both visual and tactile information to guide the behavior of an agent. We introduce Soft Fusion Contrastive Learning, an advanced version of conventional contrastive learning method and a CVAE module to utilize the alignment and complementarity within visuo-tactile representations. We demonstrate the effectiveness of our method in 12 simulated and 3 real-world environments, and our experiments show that ViTaS significantly outperforms existing baselines. Project page: https://skyrainwind.github.io/ViTaS/index.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。