从零训练机器人动作模型,实现跨场景零样本操控
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

- 自研生成与动作组件,分离预训练视觉规划和逆动力学
- 3万小时数据训练使成功率提升至44.1%,小数据集贡献显著
- 可零样本理解指令并纠正错误行为,适合复杂现实任务
世界-动作模型(WAM)通过预测未来状态指导机器人行动,支持无动作视频与带动作交互的联合学习。现有方法多依赖预训练视频生成模型,导致WAM的预训练与扩展研究不足。本文提出GE-Act 2.0,所有生成与动作组件均在操作数据上从零初始化,包含控制导向自编码器(CoAE)、单步视觉规划器(SVP)和逆动力学模型(IDM)。CoAE在高压压缩下保留动作与指令相关特征,SVP可在一次可微传递中生成完整未来状态,使视觉规划与逆动力学可分别在互补数据上预训练。随后通过知识对齐选择性优化(KASO)联合训练,仅保留行为上与实际动作兼容的预测未来,减少监督错配。在20个技能组共100个任务上评估,不进行任务微调,使用保留场景、背景、光照与物体实例。数据量从300增至3万小时,G1-OP成功率由17.1%升至44.1%,G2-90D由13.4%升至31.1%;尽管仅占训练数据2%以下,G2-90D仍提升17.7个百分点,表明跨具身迁移有效。成功覆盖19/20与18/20技能组,技能覆盖率与零样本分布外(OOD)成功率强相关(皮尔逊r=0.80,斯皮尔曼rho=0.85)。同协议下,模型在至少90%试验中正确理解对象、颜色、形状与位置指代,并能在指令冲突或常规场景关联时仍遵循指令。
原文摘要 · Abstract (English)
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。