DualDiff通过双分支扩散模型提升自动驾驶场景生成质量。
DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion
- 采用语义丰富的3D表示与跨模态融合注意力机制
- 在FID、BEV分割和3D检测上均达当前最佳表现
- 适合关注自动驾驶场景生成与多模态融合的研究者
高保真驾驶场景重建依赖于充分挖掘场景信息作为条件。然而,现有方法主要依赖3D边界框和二值地图控制前景与背景,难以捕捉场景复杂性并整合多模态信息。本文提出DualDiff,一种双分支条件扩散模型,用于增强多视角驾驶场景生成。引入占位射线采样(ORS)这一语义丰富的3D表征,结合数值化驾驶场景表征,实现对前景与背景的全面控制。为提升跨模态信息融合,设计语义融合注意力(SFA)机制,对齐并融合多模态特征。此外,提出前景感知掩码损失(FGM),以增强微小物体生成。DualDiff在FID分数上达到最先进水平,并在下游鸟瞰图(BEV)分割与3D目标检测任务中持续取得更优结果。
原文摘要 · Abstract (English)
Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information. In this paper, we propose DualDiff, a dual-branch conditional diffusion model designed to enhance multi-view driving scene generation. We introduce Occupancy Ray Sampling (ORS), a semantic-rich 3D representation, alongside numerical driving scene representation, for comprehensive foreground and background control. To improve cross-modal information integration, we propose a Semantic Fusion Attention (SFA) mechanism that aligns and fuses features across modalities. Furthermore, we design a foreground-aware masked (FGM) loss to enhance the generation of tiny objects. DualDiff achieves state-of-the-art performance in FID score, as well as consistently better results in downstream BEV segmentation and 3D object detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。