用未来视觉监督学习触觉增强的视觉语言动作模型,提升复杂操作鲁棒性。
τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

- 基于未来视觉监督在隐空间学习动作相关的时空触觉表征
- 在4个接触任务上超越现有模型,对未见物体场景具泛化能力
- 适合研究多模态机器人操作与触觉感知融合的学者
将触觉感知融入视觉-语言-动作(VLA)模型有望提升高接触交互任务的表现,但受限于任务数据稀少,如何学习有效触觉表征并适配预训练VLA模型仍具挑战。现有方法或仅关注瞬时接触状态,或使用6维力/力矩序列建模时序动态,未能充分利用高维触觉信号。为此,我们提出τ框架,借鉴联合嵌入预测架构(JEPA),从未来视觉监督中学习动作条件的时空触觉表示,并与视觉-语言特征融合生成动作。该监督仅用于训练,不增加部署开销。我们还构建了TacAura数据集,包含四个典型接触任务中的同步视觉、本体感觉和视觉触觉信号。实验表明,τ在多个任务中优于现有模型,能泛化至未见物体与场景,显著提升操作性能与鲁棒性。
原文摘要 · Abstract (English)
Incorporating tactile sensing into Vision-Language-Action (VLA) models holds promise for contact-rich manipulation, where visual observations alone often fail to capture critical cues about physical interactions. However, learning informative tactile representation while effectively adapting it to pretrained VLA models remains challenging under limited task-specific data. Existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present τ, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that τ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness. Project Page: https://cocacola-lab.github.io/tau-Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。