小米自研世界模型,兼顾场景重建与生成,提升自动驾驶仿真效果。
Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

- 用稀疏查询驱动3D重建,实现跨视角跨时间特征融合。
- 两阶段训练生成模型,4步去噪即可完成高质量视频生成。
- 集成重建与生成模块,提升生成稳定性与视觉保真度。
本文提出一个统一技术体系,解决自动驾驶世界模型的两大核心能力:世界表征与世界生成。在世界表征方面,提出WorldRec,一种由稀疏场景查询驱动的前馈重建架构。WorldRec在3D空间初始化结构化查询,通过聚合多视角、多时序特征,自然实现帧间空间一致性,生成紧凑而高保真的3D Gaussian场景表示。在世界生成方面,提出WorldGen,一种双阶段训练框架:先双向预训练,再通过三阶段渐进微调(教师强制、ODE蒸馏、动态模态分解),实现在仅4步去噪的情况下高质量在线因果视频生成。基于这两个模块,进一步构建了联合世界模型JWM,深度整合WorldRec与WorldGen,显著提升生成稳定性、帧间一致性和视觉保真度,为闭环仿真、数据合成及端到端训练提供坚实基础。
原文摘要 · Abstract (English)
This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。