arXiv:2506.17561cs.CVcs.AI2025-06NeurIPS被引 38

对比视觉与语言规划在多模态机器人中的表现,发现视觉规划更优且分层架构综合性能最强。

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

  • 构建统一架构VLA-OS,隔离网络与数据影响,专注比较规划范式与表征
  • 视觉基底规划优于语言规划,分层范式在多数任务中表现最佳
  • 适合研究复杂机器人任务规划的算法设计与可扩展性评估

近期视觉-语言-动作(VLA)模型从端到端动作生成转向先规划后执行的流程,在多种复杂长程操作任务中表现更优。然而,现有方法在网络结构、规划范式、表示方式和训练数据上差异显著,难以厘清性能提升的具体来源。为此,本文提出VLA-OS,一套支持多种规划范式的统一VLA架构,并设计涵盖刚体与柔体物体、2D与3D视觉模态、仿真与真实环境、夹持器与灵巧手等多种条件的系统性对照实验。结果表明:1)视觉基底规划表示普遍优于语言规划表示;2)分层-VLA范式在任务性能、预训练能力、泛化性、可扩展性及持续学习方面总体表现更优或相当,尽管训练与推理速度较慢。

原文摘要 · Abstract (English)

Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approaches vary significantly in terms of network architectures, planning paradigms, representations, and training data sources, making it challenging for researchers to identify the precise sources of performance gains and components to be further improved. To systematically investigate the impacts of different planning paradigms and representations isolating from network architectures and training data, in this paper, we introduce VLA-OS, a unified VLA architecture series capable of various task planning paradigms, and design a comprehensive suite of controlled experiments across diverse object categories (rigid and deformable), visual modalities (2D and 3D), environments (simulation and real-world), and end-effectors (grippers and dexterous hands). Our results demonstrate that: 1) visually grounded planning representations are generally better than language planning representations; 2) the Hierarchical-VLA paradigm generally achieves superior or comparable performance than other paradigms on task performance, pretraining, generalization ability, scalability, and continual learning ability, albeit at the cost of slower training and inference speeds.

机器人规划视觉语言动作分层决策多模态表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。