用潜在世界建模分离意图与动作,提升机器人决策稳定性。
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

- 通过VLM构建潜在世界模型,生成带意图的视觉前瞻作为中间瓶颈。
- 在RoboCasa数据集上仅需1/10演示即可达到新最好性能。
- 适合需要少样本学习和跨场景泛化的机器人控制任务。
视觉-语言-动作(VLA)模型的发展得益于预训练视觉-语言模型(VLM)。然而,现有端到端VLA大多将VLM仅作为多模态编码器,直接将视觉-语言特征映射为低层动作,未能充分发挥其高层决策潜力,且导致训练不稳定,破坏了丰富的语义表征。为此,我们提出DIAL框架,通过可微分的潜在意图瓶颈连接高层决策与低层执行。具体地,基于VLM的系统2在原始特征空间中合成潜在视觉前景,显式编码意图并作为结构瓶颈;轻量级系统1则结合预测意图与当前观测,通过潜在逆动力学解码出精确机器人动作。为保证优化稳定,采用两阶段训练:先在解耦预热阶段,系统2学习预测潜在未来,系统1在真实未来引导下学习运动控制,二者共享统一特征空间;随后无缝进入端到端联合优化。这使得动作感知梯度可控地优化VLM主干,保留预训练知识。在RoboCasa GR1桌面上基准测试中,DIAL实现新最佳性能,所需演示仅是先前方法的1/10。此外,通过利用异构人类示范,DIAL学习到物理上合理的操作先验,在真实人形机器人部署中展现出对未见物体和新配置的强零样本泛化能力。
原文摘要 · Abstract (English)
The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping vision-language features to low-level actions. This paradigm underutilizes the VLM's potential in high-level decision making and introduces training instability, frequently degrading its rich semantic representations. To address these limitations, we introduce DIAL, a framework bridging high-level decision making and low-level motor execution through a differentiable latent intent bottleneck. Specifically, a VLM-based System-2 performs latent world modeling by synthesizing latent visual foresight within the VLM's native feature space; this foresight explicitly encodes intent and serves as the structural bottleneck. A lightweight System-1 policy then decodes this predicted intent together with the current observation into precise robot actions via latent inverse dynamics. To ensure optimization stability, we employ a two-stage training paradigm: a decoupled warmup phase where System-2 learns to predict latent futures while System-1 learns motor control under ground-truth future guidance within a unified feature space, followed by seamless end-to-end joint optimization. This enables action-aware gradients to refine the VLM backbone in a controlled manner, preserving pre-trained knowledge. Extensive experiments on the RoboCasa GR1 Tabletop benchmark show that DIAL establishes a new state-of-the-art, achieving superior performance with 10x fewer demonstrations than prior methods. Furthermore, by leveraging heterogeneous human demonstrations, DIAL learns physically grounded manipulation priors and exhibits robust zero-shot generalization to unseen objects and novel configurations during real-world deployment on a humanoid robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。