arXiv:2602.06001cs.RO2026-02被引 20

融合视觉与触觉的机器人世界模型,提升接触任务中的物理理解与规划能力。

Visuo-Tactile World Models

  • 用触觉补足视觉,通过触觉推理捕捉接触物理规律。
  • 自回归推演中物体恒存性提升33%,运动合规性提高29%。
  • 零样本实机测试成功率最高提升35%,适合复杂接触任务规划。

我们提出多任务视觉-触觉世界模型(VT-WM),通过触觉推理捕捉接触过程中的物理规律。在一系列接触密集型操作任务上训练后,VT-WM显著提升了对物理世界的想象能力:在自回归推演中,物体恒存性表现提升33%,运动规律符合度提高29%。实验表明,接触动力学的具身化也促进了规划性能——在零样本真实机器人测试中,成功率最高提升35%,尤其在多步骤、高接触任务中收益显著。此外,VT-WM展现出强大下游泛化能力,仅需少量示范即可将学习到的接触动态迁移到新任务,并实现可靠规划。

原文摘要 · Abstract (English)

We introduce multi-task Visuo-Tactile World Models (VT-WM), which capture the physics of contact through touch reasoning. By complementing vision with tactile sensing, VT-WM better understands robot-object interactions in contact-rich tasks, avoiding common failure modes of vision-only models under occlusion or ambiguous contact states, such as objects disappearing, teleporting, or moving in ways that violate basic physics. Trained across a set of contact-rich manipulation tasks, VT-WM improves physical fidelity in imagination, achieving 33% better performance at maintaining object permanence and 29% better compliance with the laws of motion in autoregressive rollouts. Moreover, experiments show that grounding in contact dynamics also translates to planning. In zero-shot real-robot experiments, VT-WM achieves up to 35% higher success rates, with the largest gains in multi-step, contact-rich tasks. Finally, VT-WM demonstrates significant downstream versatility, effectively adapting its learned contact dynamics to a novel task and achieving reliable planning success with only a limited set of demonstrations.

世界模型触觉感知机器人控制多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。