arXiv:2503.08548cs.ROcs.CV2025-03被引 72

让机器人通过触觉和语言理解完成高接触任务,效果显著优于传统方法。

TLA: Tactile-Language-Action Model for Contact-Rich Manipulation

  • 用多模态语言对齐处理连续触觉反馈,生成稳健操作策略。
  • 在未见孔位和钉头形状下达成超85%成功率,显著提升动作准确率。
  • 开源2.4万组触觉指令数据,适合触觉-语言-操作研究者使用。

视觉-语言模型已取得显著进展,但针对高接触任务的语言引导机器人操作仍研究不足,尤其缺乏触觉感知的整合。为此,我们提出触觉-语言-动作(TLA)模型,通过跨模态语言对齐有效处理序列触觉反馈,实现在强接触场景下的鲁棒策略生成。同时构建了一个包含24,000对触觉动作指令的数据集,专为指尖插孔装配任务设计,为TLA训练与评估提供关键资源。实验表明,TLA在有效动作生成和动作精度上显著优于传统模仿学习方法(如扩散策略),并在未见过的装配间隙和钉头形状上实现超过85%的成功率,展现出强大的泛化能力。项目代码与数据已公开,以推动语言引导触觉操作技能学习的研究进展。

原文摘要 · Abstract (English)

Significant progress has been made in vision-language models. However, language-conditioned robotic manipulation for contact-rich tasks remains underexplored, particularly in terms of tactile sensing. To address this gap, we introduce the Tactile-Language-Action (TLA) model, which effectively processes sequential tactile feedback via cross-modal language grounding to enable robust policy generation in contact-intensive scenarios. In addition, we construct a comprehensive dataset that contains 24k pairs of tactile action instruction data, customized for fingertip peg-in-hole assembly, providing essential resources for TLA training and evaluation. Our results show that TLA significantly outperforms traditional imitation learning methods (e.g., diffusion policy) in terms of effective action generation and action accuracy, while demonstrating strong generalization capabilities by achieving over 85\% success rate on previously unseen assembly clearances and peg shapes. We publicly release all data and code in the hope of advancing research in language-conditioned tactile manipulation skill learning. Project website: https://sites.google.com/view/tactile-language-action/

触觉感知语言引导机器人操作多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。