用扩散语言模型构建视觉-语言与视觉-语言-动作模型,提升复杂规划与机器人控制性能。
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- 基于扩散语言模型构建视觉-语言模型,突破传统自回归生成的顺序限制。
- 在LIBERO等数据集上达97.2%成功率,优于π₀和GR00T-N1等主流模型。
- 适合需要并行生成与动作分块的机器人控制任务,训练收敛更快。
尽管自回归大视觉-语言模型(VLM)取得了显著进展,但其序列生成方式常限制其在复杂视觉规划与动态机器人控制中的表现。本文探索将扩散语言模型(dLLM)作为基础构建视觉-语言模型(dVLM)的潜力。我们提出Dream-VL,一个开源的扩散型视觉-语言模型,在先前dVLM中达到顶尖水平,其性能可媲美在开放数据上训练的顶级自回归模型,且在视觉规划任务中展现更优潜力。在此基础上,我们进一步推出Dream-VLA,一种基于dLLM的视觉-语言-动作模型(dVLA),通过在开放机器人数据集上持续预训练构建。我们发现扩散模型固有的双向特性天然适配动作分块与并行生成,显著加速下游微调收敛。Dream-VLA在LIBERO上实现97.2%平均成功率,于SimplerEnv-Bridge达71.4%,SimplerEnv-Fractal达60.5%,超越π₀和GR00T-N1等领先模型。同时验证了dVLM在不同训练目标下均优于自回归基线。我们已开源Dream-VL与Dream-VLA,以推动社区研究。
原文摘要 · Abstract (English)
While autoregressive Large Vision-Language Models (VLMs) have achieved remarkable success, their sequential generation often limits their efficacy in complex visual planning and dynamic robotic control. In this work, we investigate the potential of constructing Vision-Language Models upon diffusion-based large language models (dLLMs) to overcome these limitations. We introduce Dream-VL, an open diffusion-based VLM (dVLM) that achieves state-of-the-art performance among previous dVLMs. Dream-VL is comparable to top-tier AR-based VLMs trained on open data on various benchmarks but exhibits superior potential when applied to visual planning tasks. Building upon Dream-VL, we introduce Dream-VLA, a dLLM-based Vision-Language-Action model (dVLA) developed through continuous pre-training on open robotic datasets. We demonstrate that the natively bidirectional nature of this diffusion backbone serves as a superior foundation for VLA tasks, inherently suited for action chunking and parallel generation, leading to significantly faster convergence in downstream fine-tuning. Dream-VLA achieves top-tier performance of 97.2% average success rate on LIBERO, 71.4% overall average on SimplerEnv-Bridge, and 60.5% overall average on SimplerEnv-Fractal, surpassing leading models such as $π_0$ and GR00T-N1. We also validate that dVLMs surpass AR baselines on downstream tasks across different training objectives. We release both Dream-VL and Dream-VLA to facilitate further research in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。