arXiv:2609.09119cs.ROcs.AI2026-09

让机器人通过触觉与视觉协同想象,实现更精准的精细操作。

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

论文配图:DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
图 1 · 摘自论文原文
  • 用分治专家架构融合视觉、触觉与动作生成,动态调节触觉信息。
  • 在复杂操作任务中达成71%成功率和83.4%进展成功率。
  • 适合研究具身智能、多模态机器人控制的开发者与研究人员。

精细操作涉及与物理世界丰富的接触和微调交互,现有视觉-语言-动作(VLA)模型因严重视觉遮挡和复杂的接触动力学面临挑战。尽管近期工作已引入触觉感知,但多数方法仍依赖同质化多模态融合,缺乏自适应触觉整合与物理动力学显式建模。本文提出DeCAL,一种基于混合变压器(MoT)架构的物理具身精细操作模型,统一理解、想象与动作生成能力。通过接触感知门控策略实现自适应视觉-触觉融合,并提出视觉-触觉潜在协同想象机制,隐式建模物理世界动态。实验表明,DeCAL在所有任务中均达领先性能,平均成功率达71%,进展成功率为83.4%,且对未见场景具备强泛化能力。

原文摘要 · Abstract (English)

Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, lacking adaptive tactile integration and explicit modeling of physical dynamics. In this work, we present DeCAL, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination and action generation for contact-rich dexterous manipulation. Built upon a Mixture-of-Transformers (MoT) architecture, DeCAL leverages specialized experts for each capability while enabling efficient information flow among them. To effectively leverage tactile information, we introduce Adaptive Visuo-Tactile Fusion that dynamically regulates tactile interactions via a contact-aware gating strategy. Furthermore, we propose Visuo-Tactile Latent Co-Imagination to jointly model visual and tactile dynamics, equipping the policy with implicit physical world knowledge. Experimental results show that DeCAL consistently achieves state-of-the-art performance across all tasks, attaining a 71% average success rate and an 83.4% progress success rate, while also demonstrating strong generalization to unseen scenarios. The website is available at https://aureleopku.github.io/DeCAL.

具身智能触觉融合精细操作多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。