让机器人同时懂指令、会操作,还能不忘记之前学的知识。
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- 用多专家混合的训练方法,统一优化视觉理解与动作生成
- 在模拟任务中比同类模型提升33%,真实场景下表现更优
- 适合想实现智能机器人交互与自主控制的研究者
为在真实世界有效运作,机器人需融合多模态推理与精准动作生成。现有视觉-语言-动作(VLA)模型常牺牲一方能力,局限于特定任务数据,并遗忘预训练的视觉-语言知识。为此,我们提出InstructVLA,一个端到端的VLA模型,在保持大视觉-语言模型(VLM)灵活推理能力的同时,借助具身推理实现领先的操作性能。InstructVLA引入一种新训练范式——视觉-语言-动作指令微调(VLA-IT),结合多模态训练与专家混合适配,在标准VLM语料库和一个精心构建的650K样本的VLA-IT数据集上联合优化具身推理与动作生成。在域内SimplerEnv任务中,InstructVLA相比SpatialVLA提升33%。为评估泛化能力,我们提出SimplerEnv-Instruct,一个包含80个任务的基准,要求闭环控制与高层指令理解,其表现优于微调后的OpenVLA 96%,也优于使用GPT-4o辅助的动作专家29%。此外,InstructVLA在多模态任务上超越基线VLM,并通过文本推理实现推理时扩展,显著提升模拟与真实环境中的操作性能。这些结果表明InstructVLA在连接直观与可操控的人机交互与高效策略学习方面具有潜力。
原文摘要 · Abstract (English)
To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to task-specific manipulation data, and suffer catastrophic forgetting of pre-trained vision-language capabilities. To bridge this gap, we introduce InstructVLA, an end-to-end VLA model that preserves the flexible reasoning of large vision-language models (VLMs) while delivering leading manipulation performance with the help of embodied reasoning. InstructVLA introduces a novel training paradigm, Vision-Language-Action Instruction Tuning (VLA-IT), which employs multimodal training with mixture-of-experts adaptation to jointly optimize embodied reasoning and action generation on both standard VLM corpora and a curated 650K-sample VLA-IT dataset. On in-domain SimplerEnv tasks, InstructVLA achieves 33% improvement over SpatialVLA. To evaluate generalization, we introduce SimplerEnv-Instruct, an 80-task benchmark requiring closed-loop control and high-level instruction understanding, where it outperforms a fine-tuned OpenVLA by 96% and an action expert aided by GPT-4o by 29%. Additionally, InstructVLA surpasses baseline VLMs on multimodal tasks and exhibits inference-time scaling by leveraging textual reasoning to boost manipulation performance in both simulated and real-world settings. These results demonstrate InstructVLA's potential for bridging intuitive and steerable human-robot interaction with efficient policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。