让视觉语言动作模型学会用触觉理解物理交互,零样本完成复杂触控任务。
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- 融合视觉、语言、动作与触觉,通过混合位置-力控实现精准物理操作。
- 仅需少量示范即可激活模型对物理交互的先验知识,实现零样本泛化。
- 适合需要精细触控的机器人任务,如抓取、装配等接触密集型场景。
视觉语言动作(VLA)模型虽在指令理解上表现优异,但缺乏对物理交互的精确建模,尤其在高接触密度场景中难以实现精细力控。本文提出Tactile-VLA框架,深度融合视觉、语言、动作与触觉传感,引入混合位置-力控制器将模型意图转化为精确物理动作,并设计推理模块根据触觉反馈动态调整策略。实验验证其在三方面的能力:(1)实现触觉感知的指令遵循;(2)利用触觉相关的常识知识;(3)支持自适应触觉驱动的推理。关键发现:预训练的视觉语言模型已隐含对物理交互的语义理解;通过仅少量示范将其与触觉传感器连接,即可激活该先验知识,在接触密集任务中实现零样本泛化。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown remarkable achievements, driven by the rich implicit knowledge of their vision-language components. However, achieving generalist robotic agents demands precise grounding into physical interactions, especially in contact-rich scenarios where fine-grained force control is essential. We advance VLAs' implicit knowledge beyond identifying what to do, towards guiding how to physically interact with real world. This paper introduces Tactile-VLA, a novel framework that deeply fuses vision, language, action, and tactile sensing. This framework incorporates a hybrid position-force controller to translate the model's intentions into precise physical actions and a reasoning module that allows the robot to adapt its strategy based on tactile feedback. Experiments demonstrate Tactile-VLA's effectiveness and generalizability in three key aspects: (1) enabling tactile-aware instruction following, (2) utilizing tactile-relevant commonsense, and (3) facilitating adaptive tactile-involved reasoning. A key finding is that the VLM's prior knowledge already contains semantic understanding of physical interaction; by connecting it to the robot's tactile sensors with only a few demonstrations, we can activate this prior knowledge to achieve zero-shot generalization in contact-rich tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。