GigaBrain-0.7用三系统架构提升机器人通用能力,支持跨场景任务自适应。
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

- 三系统统一感知、预测与动作,提升多机器人泛化性
- 预训练超3.7万小时异构数据,实现零样本任务成功率显著提升
- 支持家庭与工业场景,适合研发通用机器人智能体的团队使用
视觉-语言-动作(VLA)模型已成为通用具身智能体的主要范式,在结构化环境中表现出强大的复杂与长程任务完成能力。然而,当前VLA系统是否能通过更优架构设计、扩展至更大更异构的数据规模,并在任务与具身形式间实现更广泛泛化,仍是未解问题。为此,我们提出GigaBrain-0.7,一种具身基础模型,显著提升了在多样化机器人形态间的泛化能力。具体而言,GigaBrain-0.7通过三系统架构统一理解、预测与动作,将预训练扩展至超过37,000小时的异构具身数据,并引入单阶段对齐训练,联合优化视觉-语言理解与多具身动作生成。相较于前代GigaBrain系列及$π_{0.5}$等前沿模型,GigaBrain-0.7在基础零样本能力、语言指令遵循及微调后任务成功率方面均有显著提升。尤其在自研Maker H01平台及主流机器人形态上,其在家庭与工业场景中均展现出强任务适应性与完成能力。所有训练代码与预训练模型权重将公开发布。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $π_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。