arXiv:2410.10429cs.CV2024-10被引 57

用扩散模型生成高保真可控的3D环境演化预测,提升自动驾驶规划能力。

DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model

  • 采用时空扩散Transformer,融合历史上下文生成高精度3D occupancy序列。
  • 在nuScenes上实现4D occupancy预测mIoU提升36.0%,长期生成效果显著。
  • 支持细粒度轨迹控制,适合需要精确环境预判的自动驾驶系统。

我们提出DOME,一种基于扩散模型的世界模型,能够根据历史占用观测预测未来的占用帧。该世界模型捕捉环境演变的能力对自动驾驶规划至关重要。相较于2D视频基世界模型,占用世界模型采用原生3D表示,具有易获取的标注和模态无关性。现有占用世界模型或因离散分词导致细节丢失,或依赖简单扩散架构,难以实现高效且可控的未来占用预测。DOME具备两大特性:(1) 高保真与长时序生成。采用时空扩散Transformer,基于历史上下文预测未来占用帧,有效捕捉时空信息,实现高保真细节与长时间序列生成。(2) 细粒度可控性。通过引入轨迹重采样方法,显著增强模型生成可控预测的能力。在广泛使用的nuScenes数据集上的大量实验表明,本方法在定性和定量评估中均优于现有基线,建立了新的最先进性能:在占用重建任务中mIoU提升10.5%、IoU提升21.2%;在4D占用预测任务中mIoU提升36.0%、IoU提升24.6%。

原文摘要 · Abstract (English)

We propose DOME, a diffusion-based world model that predicts future occupancy frames based on past occupancy observations. The ability of this world model to capture the evolution of the environment is crucial for planning in autonomous driving. Compared to 2D video-based world models, the occupancy world model utilizes a native 3D representation, which features easily obtainable annotations and is modality-agnostic. This flexibility has the potential to facilitate the development of more advanced world models. Existing occupancy world models either suffer from detail loss due to discrete tokenization or rely on simplistic diffusion architectures, leading to inefficiencies and difficulties in predicting future occupancy with controllability. Our DOME exhibits two key features:(1) High-Fidelity and Long-Duration Generation. We adopt a spatial-temporal diffusion transformer to predict future occupancy frames based on historical context. This architecture efficiently captures spatial-temporal information, enabling high-fidelity details and the ability to generate predictions over long durations. (2)Fine-grained Controllability. We address the challenge of controllability in predictions by introducing a trajectory resampling method, which significantly enhances the model's ability to generate controlled predictions. Extensive experiments on the widely used nuScenes dataset demonstrate that our method surpasses existing baselines in both qualitative and quantitative evaluations, establishing a new state-of-the-art performance on nuScenes. Specifically, our approach surpasses the baseline by 10.5% in mIoU and 21.2% in IoU for occupancy reconstruction and by 36.0% in mIoU and 24.6% in IoU for 4D occupancy forecasting.

3D生成扩散模型自动驾驶占用预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。