arXiv:2602.14979cs.RO2026-02被引 27

开源多模态机器人基础模型,打通感知、推理与规划闭环

RynnBrain: Open Embodied Foundation Models

  • 统一框架整合空间时序理解与物理推理能力
  • 在20个机器人基准上超越现有模型,30B版本表现最优
  • 适合机器人研发、具身智能研究者快速构建应用

尽管多模态基础模型进展迅速,具身智能领域仍缺乏一个统一的、基于物理世界的通用基础模型,难以实现感知、推理与规划的协同。我们提出RynnBrain,一个开源的时空基础模型,用于具身智能。该模型在统一框架下强化四大核心能力:全面的自我中心理解、多样化的时空定位、物理约束的推理以及物理感知的规划。RynnBrain包含三个规模模型(2B、8B、30B-A3B MoE)和四种针对下游任务微调的变体,分别适用于导航(RynnBrain-Nav)、规划(RynnBrain-Plan)、视觉语言动作(RynnBrain-VLA)及复杂空间推理(RynnBrain-CoP)。在20个具身智能基准和8个通用视觉理解基准上的评估表明,RynnBrain显著优于现有具身基础模型。后训练模型套件进一步验证了其两大潜力:(i)支持物理驱动的推理与规划;(ii)可作为强大预训练主干,高效适配多种具身任务。

原文摘要 · Abstract (English)

Despite rapid progress in multimodal foundation models, embodied intelligence community still lacks a unified, physically grounded foundation model that integrates perception, reasoning, and planning within real-world spatial-temporal dynamics. We introduce RynnBrain, an open-source spatiotemporal foundation model for embodied intelligence. RynnBrain strengthens four core capabilities in a unified framework: comprehensive egocentric understanding, diverse spatiotemporal localization, physically grounded reasoning, and physics-aware planning. The RynnBrain family comprises three foundation model scales (2B, 8B, and 30B-A3B MoE) and four post-trained variants tailored for downstream embodied tasks (i.e., RynnBrain-Nav, RynnBrain-Plan, and RynnBrain-VLA) or complex spatial reasoning tasks (i.e., RynnBrain-CoP). In terms of extensive evaluations on 20 embodied benchmarks and 8 general vision understanding benchmarks, our RynnBrain foundation models largely outperform existing embodied foundation models by a significant margin. The post-trained model suite further substantiates two key potentials of the RynnBrain foundation model: (i) enabling physically grounded reasoning and planning, and (ii) serving as a strong pretrained backbone that can be efficiently adapted to diverse embodied tasks.

具身智能基础模型机器人时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。