用生成模型实时模拟自动驾驶复杂场景,突破传统仿真局限。
NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

- 基于扩散模型自回归生成动作驱动的逼真视频序列
- 支持极端天气与突发行为,生成21000小时驾驶场景数据
- 可作为政策模型骨干,参数量仅为五分之一仍更优
随着自动驾驶能力提升,长尾场景下的安全评估仍是关键瓶颈。在闭环仿真中,驾驶策略模型与环境主动交互,其动作动态更新仿真器状态并直接影响下一帧传感器观测。现有基于重建的神经仿真器虽具逼真度,但受限于初始数据,难以泛化至高度动态或全新场景。为此,我们提出OmniDreams,一个从Cosmos扩散模型中微调得到的生成式世界模型,可实时自回归生成动作条件视频。借助Cosmos丰富的视觉先验及21,000小时驾驶场景的中后训练,OmniDreams能合成传统仿真难以捕捉的复杂现象,如极端天气和不可预测的动态物体行为。它通过结合历史帧、当前仿真状态和即时驾驶动作,自回归地生成逼真传感器输出。部署于Alpamayo 1策略模型与AlpaSim编排器的闭环系统中,OmniDreams作为高响应、强反应的环境,提供可扩展、全面的下一代自动驾驶策略训练与评估方案。初步结果显示,基于OmniDreams后训练的世界-动作模型(WAM)在Physical AI Autonomous Vehicles NuRec数据集上表现优于基于视觉语言模型的Alpamayo 1.5研究策略模型,且仅需其1/5参数量。这表明实时世界模型有望成为策略架构的核心基础。
原文摘要 · Abstract (English)
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations. While recent reconstruction-based neural simulators offer photorealism, they are fundamentally constrained by their initial captured data and struggle to generalize to highly dynamic or novel scenes. To overcome these limitations, we introduce OmniDreams, a foundation generative world model mid- and post-trained from the Cosmos diffusion model to autoregressively generate action-conditioned videos in real time. By leveraging the rich visual priors of Cosmos and mid- and post-training on 21k hours of driving scenarios, OmniDreams synthesizes complex, unobserved phenomena that are hard for traditional simulators to capture, such as extreme weather and unpredictable dynamic agent behaviors. Crucially, it autoregressively conditions its photorealistic sensor generation on past frames, the current simulator state, and immediate driving actions. Deployed in a closed-loop system with the Alpamayo 1 policy model and AlpaSim orchestrator, OmniDreams acts as a highly responsive, reactive environment, providing a scalable and comprehensive solution for training and evaluating next-generation autonomous driving policies. We additionally show preliminary results indicating that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters. These results highlight the potential for a real-time world model like OmniDreams to also serve as a backbone for policy architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。