arXiv:2608.27550cs.ROcs.CV2026-08被引 1

用少数据练出强通用机器人模型,靠的是更好的视觉-动作表征。

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

  • 先在多类机器人数据上预训练,保留通用视觉-语言先验,统一动作语义。
  • 在仿真、真实世界和未见机器人上均超越工业级系统,最高成功率达92.5%。
  • 仅用20%数据就超过全量数据基线,适合资源有限但追求泛化的研究者。

构建通用视觉-语言-动作(VLA)模型依赖大规模机器人数据,但实体数据采集成本高,覆盖稀疏。因此,在固定数据预算下,表征质量成为关键瓶颈:持续预训练应将有限轨迹转化为可迁移的视觉-动作知识,而非仅拟合具体动作。本文提出VLAct,一种面向VLA任务的视觉-语言模型骨干,在任务微调前基于广泛、异构、多机器人数据进行训练。它通过保留视觉-语言先验、多头连续动作协同监督及部分统一跨机器人动作布局,实现跨体感动作语义共享,同时支持任务特定动作头。在模拟、真实世界及未见机器人迁移任务中,VLAct表现稳定领先。在LIBERO-Plus和RoboTwin 2.0上,成功率分别达82.6%和92.5%,超越ABot-M0与LingBot-VLA等工业系统;在RoboDojo上,成功率排名第六,优于所有显式世界-动作模型(WAM)条目。最显著的是,在未见过的人形机器人环境RoboCasa-GR1上,仅使用20%下游轨迹的VLAct,性能超过全量数据的GR00T-N1.6基线。所有结果均基于开源数据和16卡训练规模,证明表征导向的持续预训练可在小算力下实现顶尖性能,是超越数据扩展的重要新方向。

原文摘要 · Abstract (English)

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.

机器人多模态持续预训练表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。