arXiv:2506.19850cs.CVcs.RO2025-06被引 126

统一建模视觉、语言与动作,提升机器人长时任务表现

Unified Vision-Language-Action Model

论文配图:Unified Vision-Language-Action Model
图 1 · 摘自论文原文
  • 将视觉、语言、动作视为离散序列联合建模
  • 在LIBERO上达95.5%成功率,超越前代方法10个百分点
  • 适合需要长程规划的机器人操控与自动驾驶场景

视觉-语言-动作模型(VLAs)在推进机器人操作方面备受关注。然而,以往方法主要依赖视觉-语言模型(VLM)的通用理解能力生成动作信号,常忽视视觉观测中蕴含的丰富时空与因果结构。本文提出UniVLA,一种统一且原生的多模态VLA模型,将视觉、语言和动作信号作为离散令牌序列进行自回归建模。该框架支持从大规模视频数据中灵活学习多模态任务,尤其擅长捕捉视频中的因果动态。通过后训练阶段引入世界建模,UniVLA显著提升了下游策略学习效果,尤其在长时任务中表现突出。在多个主流仿真基准(如CALVIN、LIBERO、Simplenv-Bridge)上,其性能达到新SOTA。例如,在LIBERO上平均成功率达95.5%,大幅领先pi0-FAST的85.5%。进一步实验证明其在真实世界ALOHA机器人操作与自动驾驶任务中具备广泛适用性。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language models (VLMs) to generate action signals, often overlooking the rich temporal and causal structure embedded in visual observations. In this paper, we present UniVLA, a unified and native multimodal VLA model that autoregressively models vision, language, and action signals as discrete token sequences. This formulation enables flexible multimodal tasks learning, particularly from large-scale video data. By incorporating world modeling during post-training, UniVLA captures causal dynamics from videos, facilitating effective transfer to downstream policy learning--especially for long-horizon tasks. Our approach sets new state-of-the-art results across several widely used simulation benchmarks, including CALVIN, LIBERO, and Simplenv-Bridge, significantly surpassing previous methods. For example, UniVLA achieves 95.5% average success rate on LIBERO benchmark, surpassing pi0-FAST's 85.5%. We further demonstrate its broad applicability on real-world ALOHA manipulation and autonomous driving.

机器人操控多模态建模长时任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。