arXiv:2505.09577cs.RO2025-05被引 82

视觉触觉语言融合模型提升插入手势的泛化能力

VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation

  • 融合视觉、触觉与语言输入,通过跨模态对齐生成机器人动作策略
  • 在未见过的钉子形状上达到90%以上成功率,优于扩散模型等基线
  • 采用偏好学习优化,实现从分类损失到连续控制任务的平滑过渡

尽管视觉语言模型发展迅速,但在语言驱动的机器人操作中,尤其是依赖接触的任务仍研究不足。为此,我们提出视觉-触觉-语言-动作模型(VTLA),通过跨模态语言对齐有效整合视觉与触觉输入,实现接触密集场景下的鲁棒策略生成。我们在仿真环境中构建了一个低成本多模态数据集,包含用于指尖插入任务的视觉-触觉-动作-指令对。此外,引入直接偏好优化(DPO)为VTLA模型提供类似回归的监督,有效弥合分类式下一词预测损失与连续机器人任务之间的差距。实验表明,VTLA模型在未见钉子形状上成功率超过90%,优于传统模仿学习方法(如扩散策略)和现有多模态基线(TLA/VLA)。最后,真实世界插孔实验验证了该模型出色的Sim2Real表现。

原文摘要 · Abstract (English)

While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks that extend beyond visually dominant pick-and-place scenarios. To bridge this gap, we introduce Vision-Tactile-Language-Action model, a novel framework that enables robust policy generation in contact-intensive scenarios by effectively integrating visual and tactile inputs through cross-modal language grounding. A low-cost, multi-modal dataset has been constructed in a simulation environment, containing vision-tactile-action-instruction pairs specifically designed for the fingertip insertion task. Furthermore, we introduce Direct Preference Optimization (DPO) to offer regression-like supervision for the VTLA model, effectively bridging the gap between classification-based next token prediction loss and continuous robotic tasks. Experimental results show that the VTLA model outperforms traditional imitation learning methods (e.g., diffusion policies) and existing multi-modal baselines (TLA/VLA), achieving over 90% success rates on unseen peg shapes. Finally, we conduct real-world peg-in-hole experiments to demonstrate the exceptional Sim2Real performance of the proposed VTLA model. For supplementary videos and results, please visit our project website: https://sites.google.com/view/vtla

机器人操作多模态融合偏好学习插入手势

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。