RoboBrain 2.5让机器人更懂空间与时间,实现精准3D操作与实时状态感知。
RoboBrain 2.5: Depth in Sight, Time in Mind
- 采用深度感知坐标预测,实现绝对度量约束下的3D空间推理
- 支持密集时序价值估计,可跨视角提供细粒度执行进度反馈
- 适合复杂精细操作任务的机器人系统开发与训练
我们提出RoboBrain 2.5,一个下一代具身AI基础模型,通过在高质量时空监督数据上进行大规模训练,推动通用感知、空间推理与时间建模能力。相较于前代,该模型引入两大关键升级:一是将2D像素相对定位转为深度感知坐标预测与绝对度量约束理解,实现物理约束下完整的3D操作轨迹生成(有序关键点序列);二是建立密集时序价值估计能力,可提供跨视角的步骤级进展预测与执行状态理解,为下游学习提供稳定反馈信号。两项改进共同推动具身智能向更物理化、执行感知的方向演进,适用于复杂精细操作任务。代码与模型检查点已公开于项目网站:https://superrobobrain.github.io
原文摘要 · Abstract (English)
We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on high-quality spatiotemporal supervision. Building upon its predecessor, RoboBrain 2.5 introduces two major capability upgrades. Specifically, it unlocks Precise 3D Spatial Reasoning by shifting from 2D pixel-relative grounding to depth-aware coordinate prediction and absolute metric constraint comprehension, generating complete 3D manipulation traces as ordered keypoint sequences under physical constraints. Complementing this spatial precision, the model establishes Dense Temporal Value Estimation that provides dense, step-aware progress prediction and execution state understanding across varying viewpoints, producing stable feedback signals for downstream learning. Together, these upgrades extend the framework toward more physically grounded and execution-aware embodied intelligence for complex, fine-grained manipulation. The code and checkpoints are available at project website: https://superrobobrain.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。