arXiv:2510.18313cs.CV2025-10被引 26

构建可跨状态、动作、奖励统一建模的自动驾驶世界模型

OmniNWM: Omniscient Driving Navigation World Models

  • 用联合分布建模生成多模态全景视频,实现像素级对齐
  • 通过结构化全景强制策略提升长时生成稳定性,减少偏差
  • 基于3D体素自动生成密集奖励,支持零样本轨迹控制

自动驾驶世界模型需在状态、动作和奖励三方面协同工作。现有方法常局限于单一模态建模、短时漂移及不精确的动作控制,且缺乏内在策略评估机制。本文提出OmniNWM——一种全景式全能驾驶世界模型,在统一的概率框架下解决三大维度问题。在状态建模上,生成包含RGB、语义、度量深度和3D占据的全景视频,通过联合分布建模实现多模态像素级对齐;为缓解自回归暴露偏差,提出结构化全景强制策略,利用随机流形加厚稳定长时生成。在动作建模上,引入规范几何动作编码与归一化的全景普吕克射线图,解耦运动动力学与传感器内参,实现跨异构数据集与相机配置的零样本轨迹控制。在奖励设计上,从生成的3D体积中直接推导出基于占据的密集奖励,建立可靠的闭环仿真环境以评估各类规划代理。大量实验表明,OmniNWM在生成保真度和控制精度上达到当前最优(SOTA),并在NuPlan及自有数据集上对新场景展现出显著的零样本鲁棒性。

原文摘要 · Abstract (English)

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods are typically restricted to fragmented modality modeling, short-horizon drift, and imprecise action control, while lacking intrinsic mechanisms for policy evaluation. In this paper, we introduce OmniNWM, an Omniscient panoramic Navigation World Model that addresses all three dimensions within a consistent probabilistic framework. For State, OmniNWM generates panoramic videos of RGB, semantics, metric depth, and 3D occupancy, ensuring pixel-level alignment across modalities with joint distribution modeling. To mitigate autoregressive exposure bias, we propose a structured panoramic forcing strategy to stabilize long-horizon generation via stochastic manifold thickening. For Action, we introduce canonical geometric action encoding with normalized panoramic Plücker ray-maps. This representation decouples motion dynamics from sensor intrinsics, enabling precise, zero-shot trajectory control across heterogeneous datasets and camera configurations. For Reward, we derive intrinsic occupancy-grounded dense rewards directly from generated 3D volumes, establishing a reliable closed-loop simulation cycle for evaluating diverse planning agents. Extensive experiments demonstrate that OmniNWM achieves SOTA performance in generation fidelity and control precision, with remarkable zero-shot robustness to novel scenes on NuPlan and in-house datasets with distinct camera rigs. Project page is available at https://arlo0o.github.io/OmniNWM/.

自动驾驶世界模型多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。