轻量版7B与32B模型统一视觉语言,提升机器人在真实环境中的感知推理能力。
RoboBrain 2.0 Technical Report
- 采用视觉编码器+语言模型的异构架构,支持多阶段训练。
- 32B模型在空间与时间基准上均超越现有开源及闭源模型。
- 适用于物理环境中物体交互、长期规划等实际场景,适合机器人研发者使用。
我们介绍RoboBrain 2.0,这是最新一代具身视觉-语言基础模型,旨在统一复杂物理环境中的感知、推理与规划。包含两个版本:轻量级7B模型和全尺寸32B模型,采用异构架构,融合视觉编码器与语言模型。尽管规模紧凑,但其在广泛具身推理任务中表现优异。32B版本在空间与时间基准测试中均取得领先结果,超越先前开源与专有模型。特别支持关键现实世界具身智能能力,如空间理解(例如可操作性预测、空间指代、轨迹预测)和时序决策(如闭环交互、多智能体长时程规划、场景图更新)。本报告详述模型架构、数据构建、多阶段训练策略、基础设施及实际应用。我们期望RoboBrain 2.0推动具身AI研究,并为通用具身智能体建设提供实用路径。代码、模型检查点及基准数据集可在https://superrobobrain.github.io获取。
原文摘要 · Abstract (English)
We introduce RoboBrain 2.0, our latest generation of embodied vision-language foundation models, designed to unify perception, reasoning, and planning for complex embodied tasks in physical environments. It comes in two variants: a lightweight 7B model and a full-scale 32B model, featuring a heterogeneous architecture with a vision encoder and a language model. Despite its compact size, RoboBrain 2.0 achieves strong performance across a wide spectrum of embodied reasoning tasks. On both spatial and temporal benchmarks, the 32B variant achieves leading results, surpassing prior open-source and proprietary models. In particular, it supports key real-world embodied AI capabilities, including spatial understanding (e.g., affordance prediction, spatial referring, trajectory forecasting) and temporal decision-making (e.g., closed-loop interaction, multi-agent long-horizon planning, and scene graph updating). This report details the model architecture, data construction, multi-stage training strategies, infrastructure and practical applications. We hope RoboBrain 2.0 advances embodied AI research and serves as a practical step toward building generalist embodied agents. The code, checkpoint and benchmark are available at https://superrobobrain.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。