让机器人通过触觉反馈提升操作精准度,无需重新训练基础模型。
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
- 用预训练触觉-语言模型提供语义触觉信息,辅助任务规划。
- 用扩散模型根据触觉信号优化动作,提升接触密集任务精度。
- 实测证明触觉双层级融合显著改善执行效果,适合具身智能研究者。
触觉反馈对物理世界有效交互至关重要,但当前最先进的视觉-语言-动作(VLA)模型缺乏解析和利用触觉信号的能力,限制了其在高接触任务中的表现。由于缺乏大规模多模态数据集,将触觉反馈融入系统面临挑战。我们提出VLA-Touch,一种在不微调基础VLA的前提下增强通用机器人策略的触觉感知方法。该方法包含两项关键创新:(1)利用预训练触觉-语言模型生成高层任务规划所需的语义触觉反馈;(2)采用基于扩散的控制器,结合触觉信号精细化调整VLA生成的动作,以应对接触密集型操作。通过真实世界实验验证,双重层次的触觉反馈集成显著提升了任务规划效率与执行精度。代码已开源。
原文摘要 · Abstract (English)
Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing \emph{without fine-tuning} the base VLA. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model that provides semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning efficiency while enhancing execution precision. Code is open-sourced at \href{https://github.com/jxbi1010/VLA-Touch}{this URL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。