arXiv:2607.04988cs.RO2026-07被引 10

让机器人模型同时理解语义、预测未来并生成动作,实现更强的泛化能力。

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

论文配图:InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
图 1 · 摘自论文原文
  • 基于预训练视觉语言模型,用可学习的前瞻标记在潜空间预测未来状态。
  • 在6个仿真基准上表现最优,真实世界中实现长时程执行与新指令泛化。
  • 无需从像素重建未来,保留语义理解,适合需要灵活适应的新任务场景。

统一的机器人操作模型旨在赋予单一策略以预训练视觉语言模型(VLM)的语义先验和通过未来预测学习到的物理动态。然而,现有方法常导致预训练主干语义退化,多目标间相互干扰,且在像素空间中从头学习未来预测,未利用预训练视频生成器的动力学先验。我们提出InternVLA-A1.5,基于原生VLM主干,在保持VQA与子任务预测训练的同时,附加轻量级统一专家用于连续动作生成。未来预测被重构为潜空间查询问题:少量可学习的前瞻标记在冻结的预训练视频生成模型监督下,将任务相关未来压缩为紧凑的潜码,使策略继承世界模型动力学先验,而无需学习像素级生成。推理阶段丢弃视频分支,保障实时控制。模型在120万机器人轨迹和300万多模态样本上预训练,六项仿真基准均取得最佳性能。真实世界中,保留的语义带来最强的组合泛化能力,两项设计协同支持长时程执行。

原文摘要 · Abstract (English)

Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of the pretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of a frozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1.2M robot episodes and 3M multimodal samples, InternVLA-A1.5 achieves the best overall results on all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.

机器人视觉语言模型未来预测泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。