arXiv:2602.00514cs.RO2026-02被引 1

通过视觉触觉对齐提升机器人操作的接触感知能力

Cross-Modal Visuo-Tactile Representation Learning with Action Chunking Transformers for Contact-Rich Manipulation

  • 用对比学习对齐视觉与触觉信号,构建共享表征空间
  • 触觉反馈使任务完成率从30%提升至54%
  • 适合做具身智能、多模态机器人控制的研究者

触觉反馈对高接触交互的机器人操作至关重要,但触觉信号常呈图像样态、依赖硬件且与外部视觉信息弱对齐。本文提出一种基于模仿学习的视觉-触觉对比学习框架,利用类CLIP目标将外部RGB图像与校准后的触觉图像映射到统一嵌入空间,并将其融入动作分块变换器(ACT)策略中。为此设计了一种低成本视觉-触觉夹持器(LVTG),提供可复现的数据采集平台,支持下游操作算法使用。在高接触交互任务上的实验表明,引入触觉反馈使任务平均完成率从仅依赖视觉的ACT基线30%提升至42%,而所提出的对比预训练进一步将完成率提高到54%。结果表明,显式对齐视觉与触觉观测,比直接添加未预训练的触觉图像,能为下游策略学习提供更有效的接触感知特征。

原文摘要 · Abstract (English)

Tactile feedback is important for contact-rich robotic manipulation, yet effective use of tactile observations remains challenging when tactile signals are image-like, hardware-dependent, and only weakly aligned with external visual observations. This study addresses this representation-learning problem by proposing a visuo-tactile contrastive learning framework for imitation-based manipulation. The method aligns external RGB observations and calibrated tactile images in a shared embedding space using a CLIP-style objective, and integrates the resulting representation into an Action Chunking Transformer (ACT) policy. A low-cost visuo-tactile gripper (LVTG) is proposed to provide a modular and durable sensing platform for reproducible data collection, supplying tactile observations that can be used by downstream manipulation algorithms. Experiments on contact-rich manipulation tasks show that tactile feedback improves the average task completion rate from 30% for a vision-only ACT baseline to 42%, and that the proposed contrastive pretraining further increases the completion rate to 54%. These results indicate that explicitly aligning visual and tactile observations provides more useful contact-aware features for downstream policy learning than directly adding tactile images without pretraining.

多模态学习触觉感知机器人操作对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。