一个模型通吃多种机器人任务,实现跨场景、跨机器人的统一决策。
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

- 用统一的视觉-语言-动作框架,把不同任务整合到一个模型里。
- 在多个数据集上表现优异,真实机器人测试中零样本成功率超26%。
- 支持多机器人形态,能自动适配不同机械结构和控制方式。
具身智能通常依赖于针对特定任务(如操作或导航)的专用模型,导致能力分散且泛化性差。本文探索是否可将异构的具身决策问题统一于单一视觉-语言-动作模型中。我们提出 Qwen-VLA,一个扩展自 Qwen 视觉-语言模型栈的统一具身基础模型,通过基于 DiT 的动作解码器实现连续动作与轨迹生成。Qwen-VLA 在大规模多源数据上联合预训练,涵盖机器人操作轨迹、人类第一视角示范、合成仿真数据、视觉-语言导航数据、轨迹中心监督及辅助视觉-语言数据。为支持多平台机器人,引入具身感知提示条件机制,以文本描述指定当前机器人形态与控制约定。我们将操作、导航与轨迹预测统一为动作与轨迹预测框架,实现跨机器人形态、任务类别与环境的可迁移视觉定位、空间推理与连续动作生成。在操作、导航与轨迹基准测试中,该模型展现一致的多任务性能,并在场景布局、背景、光照、物体配置及机器人形态变化下具备强分布外泛化能力。Qwen-VLA-Instruct 在 LIBERO 上达 97.9%,Simpler-WidowX 上达 73.7%,RoboTwin-Easy/Hard 上分别达 86.1%/87.2%,R2R 上达 69.0% OSR,RxR 上达 59.6% SR,真实世界 ALOHA 实验平均 OOD 成功率达 76.9%,在 DOMINO 动态操作任务上实现 26.6% 零样本成功率。
原文摘要 · Abstract (English)
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and auxiliary vision-language data. To support multiple robot platforms, we introduce embodiment-aware prompt conditioning, where robot-specific textual descriptions specify the current embodiment and control convention. We further cast manipulation, navigation, and trajectory prediction into a unified action-and-trajectory prediction framework, enabling transferable visual grounding, spatial reasoning, and continuous action generation across robot morphologies, task families, and environments. Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot embodiment. Qwen-VLA-Instruct achieves 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% OSR on R2R, 59.6% SR on RxR, 76.9% average OOD success in real-world ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。