RynnBrain 1.1是能支持机器人操作的通用具身大模型,性能超越主流开源与闭源模型。
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

- 采用统一时空与物理对齐框架训练,支持感知、空间推理与规划。
- 122B-A10B模型在多个评测中优于所有对比模型,真实机器人实验成功率更高。
- 适配多种机器人硬件,支持多任务联合训练,适合具身智能研究者使用。
我们提出RynnBrain 1.1,一个涵盖2B、9B和122B-A10B规模的具身基础模型家族。该模型基于统一的时空与物理对齐框架训练,支持具身感知、空间推理、定位与规划。相比RynnBrain 1.0,新增全模型族的接触点预测能力,并为2B和9B模型引入原生3D对齐,使表征与输出更贴合机器人操作需求。我们还构建了支持跨具身动作空间与具身特异性掩码的RynnBrain-VLA,并部署于Unitree G1、Astribot-S1与Tianji-Wuji机器人。RynnBrain 1.1在具身认知、定位与3D对齐任务中表现优异,其中122B-A10B模型在VSI-Bench、MMSI与RefSpatial-Bench上均超越所有评估的专有及开源模型。真实机器人实验表明,基于RynnBrain初始化的策略优于Qwen基线及代表性通用视觉语言模型,联合多任务与多具身训练显著提升任务得分与成功率。
原文摘要 · Abstract (English)
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。