提出一套构建强视觉语言动作模型的实用配方。
VLANeXt: Recipes for Building Strong VLA Models
- 从基础组件、感知要素、动作建模三方面系统分析设计选择。
- 在LIBERO和LIBERO-plus上超越现有方法,真实场景表现优异。
- 开源统一代码库,支持复现与新模型开发。
随着大型基础模型的兴起,视觉语言动作模型(VLAs)应运而生,利用视觉语言模型的强大视觉与语言理解能力实现通用策略学习。然而,当前VLA领域仍碎片化且探索性较强。尽管众多团队提出了各自的VLA模型,但训练协议和评估设置不一致,难以确定真正关键的设计选择。为此,我们基于统一框架和评估设置重新审视VLA设计空间。从类似RT-2的简单基线出发,系统分解三个维度的设计选择:基础组件、感知核心要素与动作建模视角。研究共提炼出12项关键发现,形成一套实用的VLA构建配方。最终成果为一个简单却高效的模型VLANeXt,其在LIBERO和LIBERO-plus基准上优于现有最先进方法,并在真实实验中表现出色。我们发布了统一且易用的代码库,支持复现结果、探索设计空间及在此基础上开发新VLA变体,代码库地址为 https://github.com/DravenALG/VLANeXt。
原文摘要 · Abstract (English)
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT-2, which is the origin of VLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. It outperforms the state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks and demonstrates strong performance in real-world experiments. We release a unified and easy-to-use codebase to reproduce our findings, explore the design space, and develop new VLA variants on top of a shared foundation. The codebase is available at https://github.com/DravenALG/VLANeXt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。