用少数据练出强通用机器人模型,靠的是更好的视觉-动作表征。
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
- 先在多类机器人数据上预训练,保留通用视觉-语言先验,统一动作语义。
- 在仿真、真实世界和未见机器人上均超越工业级系统,最高成功率达92.5%。
- 仅用20%数据就超过全量数据基线,适合资源有限但追求泛化的研究者。
构建通用视觉-语言-动作(VLA)模型依赖大规模机器人数据,但实体数据采集成本高,覆盖稀疏。因此,在固定数据预算下,表征质量成为关键瓶颈:持续预训练应将有限轨迹转化为可迁移的视觉-动作知识,而非仅拟合具体动作。本文提出VLAct,一种面向VLA任务的视觉-语言模型骨干,在任务微调前基于广泛、异构、多机器人数据进行训练。它通过保留视觉-语言先验、多头连续动作协同监督及部分统一跨机器人动作布局,实现跨体感动作语义共享,同时支持任务特定动作头。在模拟、真实世界及未见机器人迁移任务中,VLAct表现稳定领先。在LIBERO-Plus和RoboTwin 2.0上,成功率分别达82.6%和92.5%,超越ABot-M0与LingBot-VLA等工业系统;在RoboDojo上,成功率排名第六,优于所有显式世界-动作模型(WAM)条目。最显著的是,在未见过的人形机器人环境RoboCasa-GR1上,仅使用20%下游轨迹的VLAct,性能超过全量数据的GR00T-N1.6基线。所有结果均基于开源数据和16卡训练规模,证明表征导向的持续预训练可在小算力下实现顶尖性能,是超越数据扩展的重要新方向。
原文摘要 · Abstract (English)
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。