首个基于扩散模型的视觉语言动作模型,用于机器人操作
LLaDA-VLA: Vision Language Diffusion Action Models
- 用特殊动作标记替代全词汇分类,降低适配难度
- 分层解码动作序列,考虑动作间依赖关系
- 在仿真与真实机器人上均超越现有最先进方法
自回归视觉语言模型(VLM)的快速发展激发了对视觉语言动作模型(VLA)在机器人操作中的兴趣。最近,掩码扩散模型(一种不同于自回归模型的范式)在文本生成和多模态应用中展现出竞争力,催生了一系列基于扩散的视觉语言模型(d-VLM)。然而,将此类模型用于机器人策略学习仍鲜有探索。本文提出 LLaDA-VLA,首个基于预训练 d-VLM 的视觉-语言-扩散-动作模型,用于机器人操作。为有效适配 d-VLM 到机器人领域,我们引入两项关键设计:(1) 局部化特殊标记分类策略,将全词汇分类替换为特殊动作标记分类,降低适配难度;(2) 分层动作结构解码策略,分层解码动作序列,考虑动作内及跨动作的依赖关系。大量实验证明,LLaDA-VLA 在仿真与真实机器人上均显著优于现有最先进 VLA。
原文摘要 · Abstract (English)
The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked diffusion models, a paradigm distinct from autoregressive models, have begun to demonstrate competitive performance in text generation and multimodal applications, leading to the development of a series of diffusion-based VLMs (d-VLMs). However, leveraging such models for robot policy learning remains largely unexplored. In this work, we present LLaDA-VLA, the first Vision-Language-Diffusion-Action model built upon pretrained d-VLMs for robotic manipulation. To effectively adapt d-VLMs to robotic domain, we introduce two key designs: (1) a localized special-token classification strategy that replaces full-vocabulary classification with special action token classification, reducing adaptation difficulty; (2) a hierarchical action-structured decoding strategy that decodes action sequences hierarchically considering the dependencies within and across actions. Extensive experiments demonstrate that LLaDA-VLA significantly outperforms state-of-the-art VLAs on both simulation and real-world robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。