arXiv:2607.02503cs.RO2026-07被引 10

让机器人同时看和感受,精准完成高接触任务。

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

论文配图:VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
图 1 · 摘自论文原文
  • 用统一框架同步预测视觉、触觉变化和动作
  • 真实任务中平均成功率71.67%,领先现有方法35%以上
  • 特别适合需要精细触觉反馈的机械操作场景

高接触操纵需应对局部形变、压力、滑动和摩擦等线索,但这些信号在视觉中常稀疏且不可见。现有视觉-触觉策略通常直接将触觉输入用于动作预测,却很少建模动作生成过程中的触觉形变动态。本文提出VT-WAM,一种联合学习未来视觉预测、触觉形变预测与动作预测的统一流匹配模型。具体包括:(1) 异构混合变压器(Asymmetric MoT)注意力,连接首帧视觉锚点与时间序列触觉动态;(2) 接触门控的视觉-触觉-动作注意力引导(AVTAG),促使动作查询在接触阶段依赖触觉证据。在六个真实世界高接触操纵任务中,VT-WAM实现71.67%的平均成功率,较Fast-WAM提升26.67%,较OmniVTLA提升35.84%。消融实验表明,建模触觉形变动态与引导接触阶段触觉注意力均对任务表现至关重要。

原文摘要 · Abstract (English)

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies usually feed tactile observations directly into action prediction, but rarely model tactile deformation dynamics during action generation. In this paper, we introduce VT-WAM, a Visual-Tactile World Action Model that jointly learns future visual prediction, tactile deformation prediction, and action prediction within a unified flow matching framework. In particular, VT-WAM introduces (1) Asymmetric Mixture-of-Transformers (MoT) attention to bridge a first-frame visual anchor with temporal tactile dynamics, and (2) contact-gated Action-Visual-Tactile Attention Guidance (AVTAG) to encourage action queries to rely on tactile evidence during contact phases. Across six real-world contact-rich manipulation tasks, VT-WAM achieves a 71.67% average success rate, outperforming Fast-WAM by 26.67% and OmniVTLA by 35.84%. Ablations demonstrate that modeling tactile deformation dynamics and guiding contact-phase tactile attention are both important for contact-rich tasks. Project website: https://vt-wam.github.io/.

触觉感知机器人操纵多模态学习动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。