用扩散模型预测自动驾驶3D占位,更抗噪更准确
Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving

- 将3D占位预测转为生成式建模,利用扩散模型学习场景先验
- 在遮挡和低可见区域表现更优,比现有方法准确率显著提升
- 结果直接提升下游规划性能,适合实际自动驾驶系统
从视觉输入精确预测3D占位图对自动驾驶至关重要,但现有判别式方法在噪声数据、不完整观测及复杂3D结构下表现不佳。本文将3D占位预测重新建模为生成式任务,采用扩散模型学习底层数据分布并融入3D场景先验。该方法提升了预测的一致性与抗噪能力,更好处理3D空间结构的复杂性。大量实验表明,基于扩散模型的方法优于当前最优的判别式方法,在遮挡或低可见区域仍能生成更真实、准确的占位预测。此外,改进的预测显著提升了下游规划任务的表现,凸显了该方法在真实自动驾驶应用中的实际优势。
原文摘要 · Abstract (English)
Accurately predicting 3D occupancy grids from visual inputs is critical for autonomous driving, but current discriminative methods struggle with noisy data, incomplete observations, and the complex structures inherent in 3D scenes. In this work, we reframe 3D occupancy prediction as a generative modeling task using diffusion models, which learn the underlying data distribution and incorporate 3D scene priors. This approach enhances prediction consistency, noise robustness, and better handles the intricacies of 3D spatial structures. Our extensive experiments show that diffusion-based generative models outperform state-of-the-art discriminative approaches, delivering more realistic and accurate occupancy predictions, especially in occluded or low-visibility regions. Moreover, the improved predictions significantly benefit downstream planning tasks, highlighting the practical advantages of our method for real-world autonomous driving applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。