arXiv:2603.23481cs.ROcs.AI2026-03被引 12

VTAM融合视觉与触觉信号,提升复杂物理交互中的动作预测精度。

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

  • 用轻量级微调将触觉信息融入预训练视频模型,实现跨模态学习。
  • 在接触密集任务中平均成功率达90%,力觉敏感任务上比基线提升80%。
  • 无需触觉-语言配对数据,适合机器人操控、具身智能等场景。

视频-动作模型(VAMs)通过原始视频流学习隐式世界动态,生成时序一致的动作预测,已在长时序任务中表现优异。然而,在接触丰富的场景中,仅靠视觉难以准确捕捉精细的力调节和接触状态变化,导致行为不稳定或不精确。为此,我们提出视频-触觉-动作模型(VTAM),一种融合触觉感知的多模态世界建模框架。VTAM通过轻量级微调将触觉流注入预训练视频变压器,无需触觉-语言配对数据或独立触觉预训练,即可实现高效的跨模态表示学习。为稳定多模态融合,引入触觉正则化损失,平衡跨模态注意力,防止视觉表征主导。VTAM在接触密集型操作任务中表现出色,平均成功率达90%;在需要高保真力觉感知的薯片抓取任务中,相比pi 0.5基线提升80%。结果表明,触觉反馈对修正视觉估计误差至关重要,为物理具身基础模型提供了一种可扩展的解决方案。

原文摘要 · Abstract (English)

Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long-horizon tasks through visual reasoning, they remain limited in contact-rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine-grained force modulation and contact transitions are not reliably encoded in visual tokens, leading to unstable or imprecise behaviors. To bridge this gap, we introduce the Video-Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal. VTAM augments a pretrained video transformer with tactile streams via a lightweight modality transfer finetuning, enabling efficient cross-modal representation learning without tactile-language paired data or independent tactile pretraining. To stabilize multimodal fusion, we introduce a tactile regularization loss that enforces balanced cross-modal attention, preventing visual latent dominance in the action model. VTAM demonstrates superior performance in contact-rich manipulation, maintaining a robust success rate of 90 percent on average. In challenging scenarios such as potato chip pick-and-place requiring high-fidelity force awareness, VTAM outperforms the pi 0.5 baseline by 80 percent. Our findings demonstrate that integrating tactile feedback is essential for correcting visual estimation errors in world action models, providing a scalable approach to physically grounded embodied foundation models.

多模态机器人操控触觉感知动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。