arXiv:2607.11643cs.ROcs.AI2026-07被引 2

首个支持多机器人形态的高质量场景生成与可控编辑的统一模型

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

论文配图:Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
图 1 · 摘自论文原文
  • 将世界基础模型扩展至具身生成,统一处理图文生成、场景合成与视频生成
  • 在真实操作任务中将成功率从36.9%提升至63.2%,多视角一致性优异
  • 适合机器人具身智能、虚拟环境构建与可控内容生成方向的研究者

近期的基础图像与视频生成模型具备强泛化性和可控性,但其直接应用于具身场景受限于多视角一致性、几何连贯性及机器人本体约束。现有方法通常仅用少量机器人数据微调,常牺牲大规模预训练获得的视觉知识。我们提出Xiaomi-Robotics-U0,一个380亿参数的多模态自回归模型,用于统一具身生成。该模型将具身生成视为基础图像与视频生成的延伸,联合优化文本到图像生成、图像编辑、具身场景生成、具身迁移与具身视频生成。该统一框架在保留预训练世界基础模型泛化能力的同时,适配具身场景。Xiaomi-Robotics-U0是首个支持跨多机器人形态的高质量多视角场景生成,并引入结构化、可控的具身迁移以实现细粒度编辑,同时保持多视角一致性与交互动态。在单步与序列生成任务上达到顶尖水平,在人类评估中优于GPT-Image-2.0,于世界竞技场(World Arena)具身视频生成任务排名第一,将pi_0.5在复杂现实操作任务中的分布外成功率从36.9%提升至63.2%。结果表明,基础世界模型可同时作为具身世界模型与可扩展的数据引擎,助力具身智能发展。代码与检查点已公开于https://robotics.xiaomi.com/xiaomi-robotics-u0.html。

原文摘要 · Abstract (English)

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.

具身智能多视角生成可控编辑世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。