arXiv:2609.03572cs.CV2026-09

分层建模未来场景与实时决策,提升自动驾驶长程预判与响应能力。

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

论文配图:Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving
图 1 · 摘自论文原文
  • 构建慢-快双时序世界模型,分别处理长期演化与即时反应。
  • 引入动态感知潜变量,通过光流预测增强对运动变化的建模能力。
  • 联合预测下一帧图像和动作,实现观测驱动的连续决策更新。

世界模型为自动驾驶提供了前景预测与行为生成的新范式。现有方法或分离预测与决策,或在相同时间尺度上联合建模,难以兼顾长时前瞻与实时响应。本文提出Drive-HWM,一种分层慢-快世界模型框架:慢模型预测多步未来表征以捕捉场景长期演变;为显式建模驾驶环境中的丰富运动动态,引入通过光流预测学习的动态感知潜变量。基于这些未来表征,快模型采用轻量级多模态骨干网络与自回归专家,从最新观测中联合预测下一帧图像与即时动作。帧预测促使模型关注近景演化,一步动作生成则支持随新观测持续更新决策。在NAVSIM v1与v2上的大量实验表明,Drive-HWM具备优异驾驶性能;全面消融实验证实了分层设计、动态感知表征及联合预测机制的有效性。

原文摘要 · Abstract (English)

World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficult to simultaneously achieve long-horizon anticipation and responsive, observation-grounded decision making. We present Drive-HWM, a hierarchical slow--fast world modeling framework that organizes future representation prediction and action generation at complementary temporal scales. The slow world model predicts multi-step future representations to capture extended scene evolution. To explicitly model the abundant motion dynamics in driving environments, we introduce Dynamic-Aware Latents learned through optical-flow prediction. Guided by these future representations, the fast model uses a lightweight multimodal backbone and an autoregressive expert to jointly predict the next frame and the immediate action from the latest observation. Next-frame prediction encourages the fast model to capture imminent scene evolution, while one-step action generation allows decisions to be continuously updated as new observations arrive. Extensive experiments on NAVSIM v1 and v2 demonstrate the strong driving performance of Drive-HWM. Comprehensive ablation studies further validate the effectiveness of the hierarchical slow--fast design, dynamics-aware future representations, and joint next-frame and action prediction.

自动驾驶世界模型分层建模动态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。