arXiv:2511.01224cs.RO2025-11

让视觉语言动作模型高效适配多机器人协作,减少真人示范依赖。

Embodiment Transfer Learning for Vision-Language-Action Models

  • 用合成数据预训练唤醒模型,避免真实数据采集
  • 在真实机器人上实现6项任务超越OpenVLA 53.2%
  • 通过角色图思维区分多机器人功能分工,适合多机协同场景

视觉-语言-动作(VLA)模型显著推动了机器人学习的发展,支持在大规模跨体感数据上训练并微调至特定机器人。然而,当前先进的自回归VLA模型在多机器人协作中表现不佳。本文提出体感迁移学习(ET-VLA),核心为合成持续预训练(SCP),利用合成数据对模型进行新体感的预热,无需真实人类示范,大幅降低数据收集成本。SCP使模型掌握正确动作与精确动作标记数量。随后在目标体感数据上进行微调。为进一步提升多体感表现,提出具身思维图技术,将每个子任务建模为节点,帮助VLA区分各体感的功能与角色。研究以双臂机器人作为多机器人简化版本验证方法有效性。在仿真基准与三种不同双臂体感的真实机器人上均取得优异表现。实验显示,所提ET-VLA在六项真实任务中性能超越OpenVLA 53.2%。代码将开源,助力社区推进VLA模型在机器人学习中的应用。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. However, state-of-the-art autoregressive VLAs struggle with multi-robot collaboration. We introduce embodiment transfer learning, denoted as ET-VLA, a novel framework for efficient and effective transfer of pre-trained VLAs to multi-robot. ET-VLA's core is Synthetic Continued Pretraining (SCP), which uses synthetically generated data to warm up the model for the new embodiment, bypassing the need for real human demonstrations and reducing data collection costs. SCP enables the model to learn correct actions and precise action token numbers. Following SCP, the model is fine-tuned on target embodiment data. To further enhance the model performance on multi-embodiment, we present the Embodied Graph-of-Thought technique, a novel approach that formulates each sub-task as a node, that allows the VLA model to distinguish the functionalities and roles of each embodiment during task execution. Our work considers bimanual robots, a simple version of multi-robot to verify our approaches. We validate the effectiveness of our method on both simulation benchmarks and real robots covering three different bimanual embodiments. In particular, our proposed ET-VLA \space can outperform OpenVLA on six real-world tasks over 53.2%. We will open-source all codes to support the community in advancing VLA models for robot learning.

多机器人体感迁移合成数据具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。