用真实物体布局指导去噪,提升自动驾驶的鸟瞰图质量
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
- 引入扩散模型,以真实物体布局为引导净化鸟瞰图特征
- 在nuScenes上使3D检测mAP提升12.3%,NDS提升10.1%
- 无需修改模型结构,适合现有自动驾驶系统快速集成
鸟瞰图(BEV)表示在自动驾驶任务中至关重要。尽管近年来BEV生成技术取得进展,但传感器限制和学习过程带来的固有噪声仍未有效解决,导致下游任务性能下降。为此,我们提出BEVDiffuser,一种新型扩散模型,利用真实物体布局作为引导,有效对BEV特征图进行去噪。BEVDiffuser可在训练阶段以即插即用方式运行,无需修改现有BEV模型架构。在具有挑战性的nuScenes数据集上的大量实验表明,BEVDiffuser具备出色的去噪与生成能力,显著提升现有BEV模型性能:3D目标检测的mAP提升12.3%,NDS提升10.1%,且未增加额外计算开销。此外,在长尾物体检测及恶劣天气与光照条件下的显著改进,进一步验证了其在提升BEV表示质量方面的有效性。
原文摘要 · Abstract (English)
Bird's-eye-view (BEV) representations play a crucial role in autonomous driving tasks. Despite recent advancements in BEV generation, inherent noise, stemming from sensor limitations and the learning process, remains largely unaddressed, resulting in suboptimal BEV representations that adversely impact the performance of downstream tasks. To address this, we propose BEVDiffuser, a novel diffusion model that effectively denoises BEV feature maps using the ground-truth object layout as guidance. BEVDiffuser can be operated in a plug-and-play manner during training time to enhance existing BEV models without requiring any architectural modifications. Extensive experiments on the challenging nuScenes dataset demonstrate BEVDiffuser's exceptional denoising and generation capabilities, which enable significant enhancement to existing BEV models, as evidenced by notable improvements of 12.3\% in mAP and 10.1\% in NDS achieved for 3D object detection without introducing additional computational complexity. Moreover, substantial improvements in long-tail object detection and under challenging weather and lighting conditions further validate BEVDiffuser's effectiveness in denoising and enhancing BEV representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。