构建开放世界机器人数据集并提出双系统模型,提升复杂任务泛化能力。
Galaxea Open-World Dataset and G0 Dual-System VLA Model
- 用统一机器人采集真实场景行为数据,标注到子任务级别。
- 三阶段训练框架下,单体预训练显著提升长周期任务表现。
- 适合研究具身智能、多模态规划与机器人泛化能力的学者。
我们提出了Galaxea Open-World Dataset,一个大规模、多样化的机器人行为数据集,记录于真实的家庭与工作环境。所有演示均通过一致的机器人本体采集,并配有精确的子任务级语言标注,便于训练与评估。基于该数据集,我们引入G0,一种双系统框架:耦合视觉-语言模型(VLM)用于多模态规划,以及视觉-语言-动作(VLA)模型用于细粒度执行。G0采用三阶段课程训练:跨本体预训练、单本体预训练和任务特定微调。在涵盖桌面操作、少样本学习与长周期移动操作的综合性基准测试中,验证了该方法的有效性。尤其发现,单本体预训练阶段结合Galaxea数据集,在实现强性能中起关键作用。
原文摘要 · Abstract (English)
We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with precise subtask-level language annotations to facilitate both training and evaluation. Building on this dataset, we introduce G0, a dual-system framework that couples a Vision-Language Model (VLM) for multimodal planning with a Vision-Language-Action (VLA) model for fine-grained execution. G0 is trained using a three-stage curriculum: cross-embodiment pre-training, single-embodiment pre-training, and task-specific post-training. A comprehensive benchmark spanning tabletop manipulation, few-shot learning, and long-horizon mobile manipulation, demonstrates the effectiveness of our approach. In particular, we find that the single-embodiment pre-training stage, together with the Galaxea Open-World Dataset, plays a critical role in achieving strong performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。