arXiv:2510.20578cs.CVcs.RO2025-10被引 3

提出新型具身智能模型EmbodiedBrain,提升复杂任务规划能力。

EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence

  • 设计适配代理需求的数据结构与训练方法,融合步骤引导的策略优化。
  • 在多场景模拟中实现最优任务成功率,32B版本超越现有基线。
  • 开源完整数据、模型与评估体系,适合具身智能研究者使用。

实现通用人工智能(AGI)需要具备鲁棒空间感知、有效任务规划和自适应执行能力的具身智能体。然而,当前用于具身任务的大语言模型(LLMs)和多模态大语言模型(MLLMs)存在关键局限:模型设计与代理需求不匹配、实时延迟与性能难以兼顾、采用非真实、离线的评估指标。为此,我们提出EmbodiedBrain,一种具有7B和32B参数规模的新型视觉-语言基础模型。其框架采用代理对齐的数据结构,并结合大规模监督微调(SFT)与步数增强的组相对策略优化(Step-GRPO),通过引入前序步骤作为引导先验,显著提升长时序任务成功率。此外,我们构建了包含生成奖励模型(GRM)的综合性奖励系统,该模型在基础设施层面加速,提升训练效率。为实现全面验证,我们建立了涵盖通用性、规划性和端到端模拟的三部分评估体系,其中提出并开源了一个新颖且具有挑战性的仿真环境。实验结果表明,EmbodiedBrain在所有指标上均表现优异,确立了具身基础模型的新基准。为推动下一代通用型具身智能体的发展,我们已开源全部数据、模型权重及评估方法,可访问 https://zterobot.github.io/EmbodiedBrain.github.io。

原文摘要 · Abstract (English)

The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in physical environments. However, current large language models (LLMs) and multimodal LLMs (MLLMs) for embodied tasks suffer from key limitations, including a significant gap between model design and agent requirements, an unavoidable trade-off between real-time latency and performance, and the use of unauthentic, offline evaluation metrics. To address these challenges, we propose EmbodiedBrain, a novel vision-language foundation model available in both 7B and 32B parameter sizes. Our framework features an agent-aligned data structure and employs a powerful training methodology that integrates large-scale Supervised Fine-Tuning (SFT) with Step-Augumented Group Relative Policy Optimization (Step-GRPO), which boosts long-horizon task success by integrating preceding steps as Guided Precursors. Furthermore, we incorporate a comprehensive reward system, including a Generative Reward Model (GRM) accelerated at the infrastructure level, to improve training efficiency. For enable thorough validation, we establish a three-part evaluation system encompassing General, Planning, and End-to-End Simulation Benchmarks, highlighted by the proposal and open-sourcing of a novel, challenging simulation environment. Experimental results demonstrate that EmbodiedBrain achieves superior performance across all metrics, establishing a new state-of-the-art for embodied foundation models. Towards paving the way for the next generation of generalist embodied agents, we open-source all of our data, model weight, and evaluating methods, which are available at https://zterobot.github.io/EmbodiedBrain.github.io.

具身智能任务规划大模型仿真评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。