arXiv:2607.01586cs.CVcs.AI2026-07

统一框架对比视觉语言动作模型训练方法,发现语言与未来隐状态协同提升泛化能力。

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

论文配图:VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
图 1 · 摘自论文原文
  • 构建统一训练框架VLAFlow,使用相同架构和数据评估四种训练范式。
  • 联合语言监督与未来隐状态对齐的模型在多个基准上表现最稳定。
  • 适合关注机器人多模态预训练、跨任务迁移的研究者参考。

视觉-语言-动作模型(VLAs)近期推动了机器人操作的发展,但不同机器人数据预训练范式的效果难以比较,因现有模型在架构、数据、动作空间和评估协议上差异显著。本文提出VLAFlow(视觉-语言-动作流),一种统一的流匹配框架,用于可控比较VLA训练目标。基于包含约5000小时数据的异构机器人语料库OXEMix(涵盖DROID、OpenX-Embodiment、OpenX-Augmented和RoboCOIN),在相同的pi0风格架构、共享视觉语言模型(VLM)主干、动作专家及14维动作空间下,评估四种范式:仅动作建模(MindPI)、语言监督协同训练(MindLPI)、未来隐状态对齐(MindWPI)及其组合(MindLWPI)。在LIBERO、LIBERO-Plus和SimplerEnv上的实验表明,仅动作预训练对异构数据敏感;而语言监督有助于保持视觉-语言泛化能力,未来隐状态对齐则改善状态转移与动作结果建模。结合两者信号的MindLWPI在各基准上实现最稳定的迁移性能。结果提示一种元动作空间视角:语言与未来隐状态提供互补的中间约束,使异构动作监督更平滑、更具可迁移性。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in architecture, data, action space, and evaluation protocol. We present VLAFlow (Vision-Language-Action Flow), a unified flow-matching framework for controlled comparison of VLA training objectives. Using a heterogeneous robot corpus, OXEMix, containing approximately 5,000 hours of data from DROID, OpenX-Embodiment, OpenX-Augmented, and RoboCOIN, we evaluate four paradigms under the same pi0-style architecture, shared VLM backbone, action expert, and 14-dimensional action space: action-only modeling (MindPI), language-supervised co-training (MindLPI), future latent alignment (MindWPI), and their combination (MindLWPI). Experiments on LIBERO, LIBERO-Plus, and SimplerEnv show that action-only pre-training is sensitive to heterogeneous data. In contrast, language supervision helps preserve vision-language generalization, while future latent alignment improves state-transition and action-outcome modeling. By combining both signals, MindLWPI achieves the most stable overall transfer performance across benchmarks. These results suggest a meta-action space view: language and future latent representations provide complementary intermediate constraints that make heterogeneous action supervision smoother and more transferable.

机器人多模态预训练迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。