让机器人通过触觉与视觉联合预测未来,提升复杂操作的准确率。
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

- 融合视觉与触觉信号,仅在接触时激活触觉信息。
- 在6个高接触任务中平均提升动作准确率31.7%。
- 适合需要精细触觉反馈的机器人抓取与装配场景。
世界动作模型继承了世界模型的预测能力,使动作生成能基于对未来观测的预判。然而,它们主要依赖视觉,在高接触操作中常因缺乏物理交互线索而失效。本文提出Dream-Tac,一种统一的触觉-世界动作模型,联合建模动作、未来视觉观测与触觉动态。具体地,Dream-Tac引入(i)接触门控的视听融合机制,仅在接触时整合触觉信号;(ii)接触感知注意力偏置,更好地调节跨模态交互。为支持实时部署,设计双层加速策略:训练中重构接触感知偏置以保留融合路径,推理时采用基于缓存的扩散加速,实现最高2.9倍训练加速与1.8倍推理加速。在六个高接触操作任务中,Dream-Tac平均动作准确率提升31.7%,验证了统一视听触觉建模的有效性。代码已开源。
原文摘要 · Abstract (English)
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9$\times$ faster training and 1.8$\times$ faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7\% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。