用可控生成方法提升自动驾驶激光雷达数据质量
OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving
- 分阶段生成物体与场景,支持用户指令控制物体样式
- 在KITTI-360上比SOTA高17.5的FPD,语义分割提升57.47%
- 适合需要高质量3D感知训练数据的研究者和工程师
为提升复杂场景下自动驾驶的安全性,已有多种方法尝试模拟激光雷达点云数据。然而,这些方法在生成高质量、多样化且可控的前景物体方面仍面临挑战。为此,我们提出OLiDM,一种能在物体与场景层面生成高保真激光雷达数据的新框架。OLiDM包含两个核心模块:物体-场景渐进生成(OPG)模块和物体语义对齐(OSA)模块。OPG根据用户提示生成目标前景物体,并将其作为条件用于场景生成,实现物体与场景级别的可控输出,同时支持将用户定义的物体级标注关联到生成的场景中。OSA则旨在纠正前景物体与背景场景之间的错位,提升生成物体的整体质量。OLiDM在多种激光雷达生成任务及3D感知任务中均展现出广泛有效性。具体而言,在KITTI-360数据集上,其性能超越UltraLiDAR 17.5的FPD;在稀疏到密集激光雷达补全任务中,语义IoU相较LiDARGen提升57.47%。此外,主流3D检测器在mAP上提升2.4%,NDS提升1.9%,凸显其在推进物体感知类3D任务中的潜力。代码已公开于:https://yanty123.github.io/OLiDM。
原文摘要 · Abstract (English)
To enhance autonomous driving safety in complex scenarios, various methods have been proposed to simulate LiDAR point cloud data. Nevertheless, these methods often face challenges in producing high-quality, diverse, and controllable foreground objects. To address the needs of object-aware tasks in 3D perception, we introduce OLiDM, a novel framework capable of generating high-fidelity LiDAR data at both the object and the scene levels. OLiDM consists of two pivotal components: the Object-Scene Progressive Generation (OPG) module and the Object Semantic Alignment (OSA) module. OPG adapts to user-specific prompts to generate desired foreground objects, which are subsequently employed as conditions in scene generation, ensuring controllable outputs at both the object and scene levels. This also facilitates the association of user-defined object-level annotations with the generated LiDAR scenes. Moreover, OSA aims to rectify the misalignment between foreground objects and background scenes, enhancing the overall quality of the generated objects. The broad effectiveness of OLiDM is demonstrated across various LiDAR generation tasks, as well as in 3D perception tasks. Specifically, on the KITTI-360 dataset, OLiDM surpasses prior state-of-the-art methods such as UltraLiDAR by 17.5 in FPD. Additionally, in sparse-to-dense LiDAR completion, OLiDM achieves a significant improvement over LiDARGen, with a 57.47\% increase in semantic IoU. Moreover, OLiDM enhances the performance of mainstream 3D detectors by 2.4\% in mAP and 1.9\% in NDS, underscoring its potential in advancing object-aware 3D tasks. Code is available at: https://yanty123.github.io/OLiDM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。