用空间智能统一自动驾驶、机器人等多类智能体,实现通用泛化。
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
- 以空间认知为共性基础,分三步构建跨形态智能体模型。
- 在24个基准上表现媲美甚至超越现有最优模型。
- 适合研究通用智能体与跨领域迁移的学者参考。
通用具身智能需在异构智能体(如自动驾驶、机器人、无人机)间实现强泛化能力。然而,现有方法在训练统一模型时常面临长尾数据、梯度干扰和灾难性遗忘等问题,难以兼顾通用泛化与领域专精。本文提出ACE-Brain-0,一个将空间推理、自动驾驶与具身操作统一于单一多模态大语言模型的通用基底大脑。核心洞察是:尽管车辆、机器人和无人机形态差异显著,但均需建模三维心理空间,因此空间智能可作为跨形态通用支撑。基于此,我们提出Scaffold-Specialize-Reconcile(SSR)范式——先建立共享空间基础,再培养领域专家,最后通过无数据模型融合协调。同时采用Group Relative Policy Optimization(GRPO)增强综合能力。大量实验表明,ACE-Brain-0在24个空间与具身相关基准上达到竞争力甚至领先水平。
原文摘要 · Abstract (English)
Universal embodied intelligence demands robust generalization across heterogeneous embodiments, such as autonomous driving, robotics, and unmanned aerial vehicles (UAVs). However, existing embodied brain in training a unified model over diverse embodiments frequently triggers long-tail data, gradient interference, and catastrophic forgetting, making it notoriously difficult to balance universal generalization with domain-specific proficiency. In this report, we introduce ACE-Brain-0, a generalist foundation brain that unifies spatial reasoning, autonomous driving, and embodied manipulation within a single multimodal large language model~(MLLM). Our key insight is that spatial intelligence serves as a universal scaffold across diverse physical embodiments: although vehicles, robots, and UAVs differ drastically in morphology, they share a common need for modeling 3D mental space, making spatial cognition a natural, domain-agnostic foundation for cross-embodiment transfer. Building on this insight, we propose the Scaffold-Specialize-Reconcile~(SSR) paradigm, which first establishes a shared spatial foundation, then cultivates domain-specialized experts, and finally harmonizes them through data-free model merging. Furthermore, we adopt Group Relative Policy Optimization~(GRPO) to strengthen the model's comprehensive capability. Extensive experiments demonstrate that ACE-Brain-0 achieves competitive and even state-of-the-art performance across 24 spatial and embodiment-related benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。