8B参数模型实现物理智能闭环,通用机器人任务表现领先。
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

- 构建统一架构的具身基础模型,支持规划、纠错与指物等综合能力。
- 在24个具身视觉语言基准中16项达最优,仅用80亿参数超越主流模型。
- 可少量微调成视觉语言动作模型,真实机器人实测展现强泛化能力。
我们提出Embodied-R1.5,一种统一的具身基础模型(EFM),整合了具身认知、任务规划、纠错与指物等能力,旨在实现通用物理智能。通过三个自动化数据构建管道显著扩展关键能力的数据覆盖,构建超过150亿词元的大规模数据系统,并设计多任务平衡强化学习方案以缓解异构任务冲突。进一步提出规划-定位-纠错(PGC)闭环框架,使单一模型能自主执行并自我修正长时程任务。仅含80亿参数的Embodied-R1.5在24个具身视觉语言模型基准中取得16项最佳性能,超越Gemini-Robotics-ER-1.5和GPT-5.4等领先模型。得益于内嵌的具身能力,该模型仅需少量数据即可微调为视觉语言动作模型(VLA),在4个主流操作基准套件上超越π₀.₅。我们还开展大量零样本真实机器人实验,验证其在指令遵循、可操作性识别、关节物体操控及复杂长程任务中的表现,展示出对物理世界的强大泛化能力。我们开源模型权重、数据集、训练代码及专用于具身任务评估的EmbodiedEvalKit,推动未来具身基础模型研究。
原文摘要 · Abstract (English)
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $π_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。