arXiv:2606.29089cs.RO2026-06

用视觉增强方式让视觉语言动作模型感知触觉,提升复杂操作成功率。

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models

论文配图:TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models
图 1 · 摘自论文原文
  • 将触觉剪切力场转为视觉向量叠加在图像上,不改模型架构。
  • 在4个高接触任务中成功率达78%,远超仅用视觉的50%以下表现。
  • 无需触觉预训练,计算开销小,适合实际机器人部署。

视觉-语言-动作(VLA)模型通过大规模视觉与语言预训练,在视觉、语义和空间任务变化中展现出强大推理能力。然而,它们对接触力仍缺乏感知,而这些力虽在视觉反馈中不明显,却是高接触操作的核心。触觉传感可直接测量力,但将其融入VLA困难:触觉数据未出现在用于预训练的大规模语料中,作为新模态加入会引发分布偏移,削弱预训练带来的优势。我们提出触觉标注提示(TAP-VLA),一种仅通过视觉增强而非结构改变来提供触觉反馈的简单框架。TAP-VLA从视觉-触觉传感器提取剪切力场,并将其作为空间定位向量叠加到策略已使用的多视角RGB图像上,生成清晰可解释的触觉提示,且保持在VLA的原生观测空间内。由于架构未变,该方法无需触觉预训练,计算开销极小,且贴近预训练分布。在四个高接触任务中,TAP-VLA在78%的试验中成功,显著优于仅用视觉微调的不足50%以及其它触觉融合基线——包括部分基线表现仅略高于随机猜测。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate impressive reasoning over visual, semantic, and spatial task variations by leveraging large-scale vision and language pre-training. They remain, however, largely blind to contact forces, which seldom manifest clearly in visual feedback but are central to contact-rich manipulation. Tactile sensing measures these forces directly, but integrating it into VLAs is difficult: tactile data is absent from the large-scale corpora used to pre-train VLAs, so adding it as a new input modality induces a distribution shift that erodes the very pre-training that makes VLAs effective. We propose Tactile Annotation Prompting for Vision-Language-Action models (TAP-VLA), a simple framework that supplies tactile feedback through visual augmentation rather than architectural change. TAP-VLA extracts shear fields from visuo-tactile sensors and overlays them as spatially-grounded vectors onto the multi-view RGB images the policy already consumes, yielding a clear, interpretable tactile cue in the VLA's native observation space. Because the architecture is untouched, the approach requires no tactile pre-training, adds negligible compute, and stays close to the pre-training distribution. Across four contact-rich tasks, TAP-VLA succeeds on 78% of trials, compared to under 50% for vision-only fine-tuning and alternative tactile-fusion baselines -- including tasks where the baselines perform no better than chance.

触觉感知视觉增强机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。