让自动驾驶模型生成更符合真实4D场景的未来预测。
4D-WAM: 4D Consistent World Modeling for Autonomous Driving

- 用几何基础模型监督未来帧,强制保持4D一致性。
- 在早期高噪声阶段强化训练,提升决策准确性。
- 适合需要可靠场景预测的自动驾驶系统研究者。
新兴的世界-动作模型(WAMs)通过联合建模未来驾驶场景演化与轨迹规划,在自动驾驶中展现出良好性能。然而,现有WAMs通常基于2D视频数据训练,无法理解底层4D场景结构,导致生成的未来预测虽视觉合理却存在4D不一致,误导下游规划。为此,我们提出4D-WAM,利用几何基础模型在训练时提供4D感知监督:将模型预测的未来帧输入几何基础模型,根据其4D感知响应构建一致性损失。该损失促使模型在训练中理解、表征并预测物理上一致的4D场景,且不增加推理开销。此外,我们发现WAMs存在早期决策现象,提出面向决策的采样策略,重点强化在早期高噪声阶段的监督,此时驾驶决策主要形成。通过将4D监督传递至这一关键阶段,策略进一步提升轨迹规划效果。大量实验表明,4D-WAM有效建模4D一致的场景演化,在挑战性的NAVSIM-v1和NAVSIM-v2基准上达到当前最优性能。
原文摘要 · Abstract (English)
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visually plausible yet 4D inconsistent future predictions that mislead downstream planning. To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Specifically, we feed WAM-predicted future frames into a geometric foundation model, and use 4D-aware responses to define a 4D consistency loss. This loss encourages the model to understand, represent, and predict physically consistent 4D scenes during training, without additional inference cost. Moreover, we identify an early-decision phenomenon in WAMs and propose a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages, where driving decisions are primarily formed. By propagating 4D supervision to this critical decision-formation phase, the proposed strategy further improves trajectory planning. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。