arXiv:2603.24581cs.CVcs.RO2026-03被引 9

用压缩的隐空间建模世界状态,实现高效端到端自动驾驶规划。

Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving

  • 通过可学习查询将多视角图像压缩为紧凑场景令牌,提升空间感知能力。
  • 采用因果Transformer预测未来世界状态,在NAVSIM v2上达89.3 EPDMS新高。
  • 模型仅104M参数,训练数据少却超越现有无感知方法,适合资源受限场景。

我们提出Latent-WAM,一种高效的端到端自动驾驶框架,通过空间感知与动态信息丰富的隐空间世界表征实现优异轨迹规划。现有基于世界模型的规划器存在表征压缩不足、空间理解有限、时序动态利用不充分等问题,导致在数据与算力受限条件下规划效果不佳。Latent-WAM引入两个核心模块:空间感知压缩世界编码器(SCWE),从基础模型中提取几何知识,通过可学习查询将多视角图像压缩为紧凑场景令牌;动态隐空间世界模型(DLWM),采用因果Transformer自回归预测未来世界状态,条件于历史视觉与运动表征。在NAV SIM v2和HUG SIM上的大量实验表明,该框架达到新最佳性能:NAV SIM v2上89.3 EPDMS,HUG SIM上28.9 HD-Score,相比最优前序无感知方法提升3.2 EPDMS,且显著减少训练数据需求,模型规模仅为104M参数。

原文摘要 · Abstract (English)

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-based planners suffer from inadequately compressed representations, limited spatial understanding, and underutilized temporal dynamics, resulting in sub-optimal planning under constrained data and compute budgets. Latent-WAM addresses these limitations with two core modules: a Spatial-Aware Compressive World Encoder (SCWE) that distills geometric knowledge from a foundation model and compresses multi-view images into compact scene tokens via learnable queries, and a Dynamic Latent World Model (DLWM) that employs a causal Transformer to autoregressively predict future world status conditioned on historical visual and motion representations. Extensive experiments on NAVSIM v2 and HUGSIM demonstrate new state-of-the-art results: 89.3 EPDMS on NAVSIM v2 and 28.9 HD-Score on HUGSIM, surpassing the best prior perception-free method by 3.2 EPDMS with significantly less training data and a compact 104M-parameter model.

自动驾驶隐空间建模端到端轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。