arXiv:2601.02456cs.RO2026-01被引 43

融合视觉、语言与动作的机器人模型,提升动态任务表现。

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

  • 采用统一混合变压器架构,协同理解、预测与执行
  • 在12项真实任务中,动态操作性能提升26.7%
  • 适合需要感知与动作联动的机器人研究者

现有视觉-语言-动作(VLA)模型多基于多模态大语言模型,虽具备优秀语义理解能力,但缺乏对物理世界动态的推断。近期方法转向世界模型,通过视频预测实现动态建模,但常因语义缺失和预测误差导致鲁棒性差。为此,我们提出InternVLA-A1,采用统一的混合变压器架构,协调场景理解、视觉前瞻生成与动作执行三个专家模块,通过统一掩码自注意力机制实现无缝交互。基于InternVL3与Qwen3-VL,我们构建了2B和3B参数规模的模型,在超过6.92亿帧的真实机器人数据、仿真数据与人类视频数据上进行预训练。该混合训练策略有效利用仿真数据多样性,同时减小仿真到现实的差距。在12项真实机器人任务与RoboTwin 2.0仿真基准上的评估显示,相比pi0.5,InternVLA-A1在静态操作任务上提升4.4%,在模拟基准上提升2.6%,在动态操作任务上提升26.7%。

原文摘要 · Abstract (English)

Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, but they inherently lack the capability to deduce physical world dynamics. Consequently, recent approaches have shifted toward World Models, typically formulated via video prediction; however, these methods often suffer from a lack of semantic grounding and exhibit brittleness in the presence of video prediction errors. To synergize semantic understanding with dynamic predictive capabilities, we present InternVLA-A1. This model employs a unified Mixture-of-Transformers architecture, coordinating three experts for scene understanding, visual foresight generation, and action execution. These components interact seamlessly through a unified masked self attention mechanism. Building upon InternVL3 and Qwen3-VL, we instantiate InternVLA-A1 at 2B and 3B parameter scales. We pre-train these models on heterogeneous data sources over real-world robot data, synthetic simulation data, and human videos, covering over 692M frames. This hybrid training strategy effectively harnesses the diversity of synthetic simulation data while minimizing the sim-to-real gap. We evaluated InternVLA-A1 on 12 real-world robotic tasks and a simulation benchmark. The results show that InternVLA-A1 consistently outperforms prior leading models: compared with pi0.5, it achieves +4.4\% on static manipulation tasks and +2.6\% on the RoboTwin 2.0 simulation benchmark, and delivers a +26.7\% boost on dynamic manipulation tasks.

机器人操控多模态模型视觉预测动态任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。