arXiv:2607.14187cs.AIcs.RO2026-07被引 1

让机器人同时理解语言和视觉,生成有物理意义的行动计划。

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

论文配图:RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
图 1 · 摘自论文原文
  • 用统一模型融合语言与视觉想象,共同规划任务步骤。
  • 在真实机器人上无需大量动作数据预训练即表现良好。
  • 适合研究具身智能、多模态决策与机器人规划的团队。

具身认知要求智能体将高层任务推理与目标物理状态相连接。我们提出 Hy-Embodied-RxBrain,一个具备联合语言-视觉推理与想象能力的具身认知基础模型。不同于侧重场景理解或文本决策的视觉-语言模型,或仅预测未来视觉状态的生成世界模型,RxBrain 在单一规划序列中表示具身计划,其中语言提供计划的抽象结构(如任务分解、规划原语、约束、时序顺序与决策逻辑),而视觉想象通过世界状态预测与联合子目标规划来具体化该结构,关联每个步骤与中间及最终物理状态。模型采用统一的多模态 Transformer 混合架构,支持语言、图像与视频的理解与生成。为训练该能力,我们构建自动管道,将具身视频转化为包含文本-视觉联合规划监督的标注数据,通过分解视频为规划步骤并对其视觉状态转移进行对齐。我们进一步提出 RxBrain-Bench,评估模型能否通过联合文本与视觉组件表征具身计划,而非独立理解或生成。实验表明,RxBrain保持了具身理解与生成能力,能生成耦合文本推理、世界状态预测与联合子目标规划的计划。我们还将模型扩展至连续机器人动作生成,在真实机器人上表现出色,且无需大规模动作数据预训练。这些结果为具身认知基础模型的发展提供了初步探索。

原文摘要 · Abstract (English)

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.

具身智能多模态机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。