arXiv:2410.00337cs.CV2024-10被引 18

用3D语义平面图控制生成逼真街景,解决自动驾驶数据标注难题

SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs

  • 用3D语义多平面图(MPIs)作为条件输入,让2D扩散模型理解3D几何信息
  • 在nuScenes数据集上生成的多视角图像与给定占用标签高度一致,保真度高
  • 适合用于感知模型训练和仿真,可无限生成带标注的可控街景数据

自动驾驶发展越来越依赖高质量标注数据,尤其在3D占用预测任务中,占用标签需密集3D标注,人工成本高昂。本文提出SyntheOcc,一种基于扩散模型的图像生成方法,通过驾驶场景中的占用标签条件,生成逼真且几何可控的街景图像。该方法解决了如何高效将3D几何信息编码为2D扩散模型条件输入的关键挑战。创新性地引入3D语义多平面图像(MPIs),提供全面且空间对齐的3D场景描述作为条件。结果表明,SyntheOcc能生成与给定几何标签一致的逼真多视角图像和视频。在nuScenes数据集上的定性和定量评估验证了其生成可控占用数据集的有效性,可作为感知模型的数据增强手段。

原文摘要 · Abstract (English)

The advancement of autonomous driving is increasingly reliant on high-quality annotated datasets, especially in the task of 3D occupancy prediction, where the occupancy labels require dense 3D annotation with significant human effort. In this paper, we propose SyntheOcc, which denotes a diffusion model that Synthesize photorealistic and geometric-controlled images by conditioning Occupancy labels in driving scenarios. This yields an unlimited amount of diverse, annotated, and controllable datasets for applications like training perception models and simulation. SyntheOcc addresses the critical challenge of how to efficiently encode 3D geometric information as conditional input to a 2D diffusion model. Our approach innovatively incorporates 3D semantic multi-plane images (MPIs) to provide comprehensive and spatially aligned 3D scene descriptions for conditioning. As a result, SyntheOcc can generate photorealistic multi-view images and videos that faithfully align with the given geometric labels (semantics in 3D voxel space). Extensive qualitative and quantitative evaluations of SyntheOcc on the nuScenes dataset prove its effectiveness in generating controllable occupancy datasets that serve as an effective data augmentation to perception models.

3D生成自动驾驶扩散模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。