arXiv:2603.25406cs.RO2026-03中稿 · ACM MM 2026被引 8

用统一离散扩散模型实现视觉语言动作的端到端生成,提升长序列任务一致性。

MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation

  • 采用共享离散标记空间,同步去噪目标图像与动作片段。
  • 在LIBERO上达98.0%成功率,CALVIN上成功序列长度达4.78。
  • 无需额外世界模型,适合复杂长程机器人任务部署。

视觉-语言-动作(VLA)模型将视觉观测和自然语言指令映射为机器人动作;然而,分层与自回归范式常带来架构开销,累积长时序误差,并需辅助模块捕捉环境动态。为此,我们提出MMaDA-VLA,一种完全原生的预训练离散扩散VLA模型,实现了多模态理解与生成的统一。具体而言,MMaDA-VLA使用共享的离散标记空间,联合去噪未来目标观测与动作块,使动作基于预测的视觉结果,无需辅助世界模型。该方法实现并行、无序的优化,提升长时序一致性。大量实验与全面分析表明,MMaDA-VLA在LIBERO上平均成功率达98.0%,在CALVIN上平均成功序列长度为4.78,且在真实场景中表现优异。项目页面见https://yliu-cs.github.io/MMaDA-VLA。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actions; however, hierarchical and autoregressive paradigms often incur architectural overhead, accumulate long-horizon errors, and require auxiliary modules to capture environment dynamics. To this end, we present MMaDA-VLA, a fully native, pretrained discrete diffusion VLA that unifies multi-modal understanding and generation. Specifically, MMaDA-VLA uses a shared discrete token space to jointly denoise a future goal observation and an action chunk, grounding actions in predicted visual outcomes without an auxiliary world model. In this way, parallel, order-free refinement improves long-horizon consistency. Extensive experiments and comprehensive analyses demonstrate that MMaDA-VLA achieves an average success rate of 98.0\% on LIBERO and an average successful sequence length of 4.78 on CALVIN, while performing strongly in real-world settings. The project page is available at https://yliu-cs.github.io/MMaDA-VLA.

机器人扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。