构建可自主规划、执行与改进的通用机器人智能体,统一感知、决策与学习闭环。
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

- 五功能一体化:空间感知、决策、交互、自监控与自进化协同工作
- 80亿参数模型在15个基准上14项超越前代,支持跨任务泛化
- 适用于需要持续学习与环境适应的机器人系统开发
具身智能正从孤立的感知或动作模块迈向能理解目标、规划行动、通过机器人身体执行、监控进展并从经验中改进的物理智能体。现有系统仅部分解决该闭环:端到端策略虽生成动作但缺乏空间推理与执行评估,而机器人-代理系统虽协调工具却无法共享表征。为此,我们提出ACE-Brain-0.5,一个统一的具身基础模型,将机器人智能组织为五个耦合功能:空间感知、决策、具身交互、自我监控与自我改进。基于已建立空间智能共享架构的ACE-Brain-0,ACE-Brain-0.5将以理解为中心的模型扩展为闭环基础模型。单个80亿参数骨干网络实现前四项功能:定位物体与可用性,推理三维与自我中心空间关系,将指令分解为子目标,生成导航与操作动作,并估计进度用于验证与恢复。为避免跨任务干扰,引入SSR+,在任务向量融合后增加重激活阶段。第五项自改进功能由配套框架实现,通过运行结果更新外部执行状态,包括任务模式、空间记忆与失败恢复案例。在十五个基准测试中,ACE-Brain-0.5在18项空间感知与定位任务中的14项优于ACE-Brain-0,导航与操作性能具有竞争力,并在已知与未知场景中均表现出强进度估计能力。这些成果标志着迈向通用物理智能体的重要一步。
原文摘要 · Abstract (English)
Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve from experience. Existing systems address this loop only in parts: end-to-end policies generate actions but often lack spatial reasoning, planning, and execution assessment, while robot-agent systems orchestrate tools or specialists but do not learn a shared representation. This fragmentation limits general Physical Agentic AI. We present ACE-Brain-0.5, a unified embodied foundation model that organizes robot intelligence into five coupled functions: spatial perception, decision making, embodied interaction, self-monitoring, and self-improvement. Built on ACE-Brain-0, which established spatial intelligence as a shared scaffold across robot platforms, ACE-Brain-0.5 extends an understanding-centric model into a closed-loop foundation model. A single 8B backbone instantiates the first four functions: grounding objects and affordances, reasoning over 3D and egocentric spatial relations, decomposing instructions into subgoals, generating navigation and manipulation actions, and estimating progress for verification and recovery. To unify these capabilities without cross-task interference, we introduce SSR+, which extends Scaffold-Specialize-Reconcile with a Reactivate stage after task-vector merging. The fifth function, self-improvement, is realized by a companion framework that updates external execution state, including task schemas, spatial memory, and failure-recovery cases, from rollouts. Across fifteen benchmarks, ACE-Brain-0.5 improves over ACE-Brain-0 on 14 of 18 spatial perception and grounding benchmarks, achieves competitive navigation and manipulation performance, and provides strong progress estimation in ID and OOD settings. Together, these results mark an early step toward general Physical Agentic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。