arXiv:2501.18867cs.CVcs.AI2025-01ICML被引 95

统一理解与预测的视觉语言动作模型,提升机器人对空间细节的感知能力。

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

  • 融合多模态理解与未来预测任务,兼顾高层语义与低层空间信息
  • 在Calvin ABC-D基准上比现有最优方法提升33%成功率
  • 适合需要精准空间操作的机器人控制任务

视觉-语言-动作(VLA)模型近年来借助预训练视觉-语言模型(VLM)提升了泛化能力。这类模型通常在视觉-语言理解任务上预训练,具备丰富的语义知识与推理能力,但现有研究发现它们往往关注高层语义而忽略低层特征,难以捕捉细致的空间信息和物理动态,而这对于具身控制任务至关重要。本文探讨了VLA的训练范式,提出 extbf{UP-VLA}——一种同时包含多模态理解与未来预测目标的统一训练框架,以增强高阶语义理解和低阶空间感知。实验表明,UP-VLA在Calvin ABC-D基准上相比先前最先进方法实现33%的性能提升,并在真实世界操作任务中表现出更高的成功率,尤其在依赖精确空间信息的任务中优势显著。

原文摘要 · Abstract (English)

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich semantic knowledge and reasoning abilities. However, prior research has shown that VLMs often focus on high-level semantic content and neglect low-level features, limiting their ability to capture detailed spatial information and understand physical dynamics. These aspects, which are crucial for embodied control tasks, remain underexplored in existing pre-training paradigms. In this paper, we investigate the training paradigm for VLAs, and introduce \textbf{UP-VLA}, a \textbf{U}nified VLA model training with both multi-modal \textbf{U}nderstanding and future \textbf{P}rediction objectives, enhancing both high-level semantic comprehension and low-level spatial understanding. Experimental results show that UP-VLA achieves a 33% improvement on the Calvin ABC-D benchmark compared to the previous state-of-the-art method. Additionally, UP-VLA demonstrates improved success rates in real-world manipulation tasks, particularly those requiring precise spatial information.

具身智能视觉语言机器人控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。