arXiv:2501.14729cs.CV2025-01ICCV被引 60

统一建模驾驶场景理解与生成,提升自动驾驶预测能力

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

  • 用鸟瞰图融合多视角信息,结合大模型因果注意力注入世界知识
  • 在nuScenes和OmniDrive数据集上生成误差降低32.4%,理解指标提升8.0%
  • 适合自动驾驶感知与规划方向研究者,尤其关注场景生成与理解融合

驾驶世界模型(DWMs)已成为自动驾驶中未来场景预测的关键。然而现有模型仅聚焦场景生成,缺乏对驾驶环境的解析与推理能力。本文提出统一的驾驶世界模型HERMES,通过统一框架实现三维场景理解与未来场景演化(生成)的协同。HERMES采用鸟瞰图(BEV)表示,整合多视角空间信息并保留几何关系与交互。引入世界查询机制,借助大语言模型中的因果注意力将世界知识注入BEV特征,增强理解与生成任务的上下文表达。在nuScenes和OmniDrive-nuScenes数据集上的实验证明,该方法达到当前最优性能:生成误差降低32.4%,理解指标CIDEr提升8.0%。模型与代码将公开于https://github.com/LMD0311/HERMES。

原文摘要 · Abstract (English)

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we present a unified Driving World Model named HERMES. We seamlessly integrate 3D scene understanding and future scene evolution (generation) through a unified framework in driving scenarios. Specifically, HERMES leverages a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information while preserving geometric relationships and interactions. We also introduce world queries, which incorporate world knowledge into BEV features via causal attention in the Large Language Model, enabling contextual enrichment for understanding and generation tasks. We conduct comprehensive studies on nuScenes and OmniDrive-nuScenes datasets to validate the effectiveness of our method. HERMES achieves state-of-the-art performance, reducing generation error by 32.4% and improving understanding metrics such as CIDEr by 8.0%. The model and code will be publicly released at https://github.com/LMD0311/HERMES.

自动驾驶世界模型场景生成大模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。