arXiv:2606.08520cs.RO2026-06

用机器人轨迹数据渐进式训练,让视觉语言模型学会控制机器人。

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data

论文配图:Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data
图 1 · 摘自论文原文
  • 用机器人场景和动作轨迹构建视觉语言数据,逐步衔接预训练与控制任务。
  • 三阶段训练提升模型在真实机器人上的泛化能力,仅需少量演示即可适应新环境。
  • 适合研究具身智能、机器人控制与多模态模型迁移的开发者参考。

视觉语言模型(VLM)是强大的通用推理工具,但将其转化为机器人控制策略(VLAs)却十分困难。根本原因在于双重差距:VLM 在互联网级图像上训练,以理解语言为目标;而VLAs 需要感知机器人场景并预测动作。直接在机器人动作数据上微调VLM,迫使模型同时跨越两个鸿沟,学习曲线陡峭,预训练中获得的泛化能力往往退化而非迁移。我们提出通过合适的中间数据逐步弥合这一差距。引入‘具身轨迹耦合(ETC)’数据——源自同一机器人场景和轨迹的视觉-语言监督信号。由于ETC数据共享机器人操作的视觉语境,同时保留熟悉的语言理解目标,因此成为从VLM预训练到VLA微调的自然过渡。基于此,设计三阶段训练流程:分布桥接首先使VLM适应具身视觉语言语义;目标桥接逐步将模型转向动作预测,同时保留已有表征;保持性适配最终将策略专用于目标部署领域。进一步发现,将任务相关的域外ETC数据与少量动作数据混合,可使模型在不依赖额外机器人演示的情况下,泛化至新的视觉-语言条件。仿真与真实机器人实验表明,这种渐进式桥接策略是将VLM泛化能力成功迁移为鲁棒可部署机器人策略的关键。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are powerful general-purpose reasoners, yet converting them into robot control policies (VLAs) is surprisingly difficult. The root cause is a two-fold gap: VLMs are trained on internet-scale images with language-understanding objectives, while VLAs must perceive robot scenes and predict motor actions. Fine-tuning a VLM directly on robot action data forces the model to cross both gaps at once -- the learning curve is steep and the rich generalizations learned during pretraining tend to degrade rather than transfer. We argue that this gap can be bridged gradually with the right intermediate data. We introduce \emph{embodied trajectory-coupled (ETC) data} -- vision-language supervision derived from the same robot scenes and trajectories used for action learning. Because ETC data shares the visual context of robot operation while retaining familiar language-understanding objectives, it provides a natural stepping stone between VLM pretraining and VLA fine-tuning. Building on this, we design a three-stage training recipe. Distribution Bridging first adapts the VLM to embodied visual-language semantics. Objective Bridging then gradually shifts the model toward action prediction while preserving the acquired representations. Retentive Adaptation finally specializes the policy to the target deployment domain. We further show that mixing task-relevant out-of-distribution ETC data with a small amount of action data enables the model to generalize to novel visual-language conditions without requiring additional robot demonstrations. Simulation and real-robot experiments confirm that this gradual bridging strategy is the key to transferring VLM generalization into robust, deployable robot policies.

机器人控制多模态迁移学习具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。