用高斯点云重建场景,精准放置3D物体提升检测效果
Gaussian Splatting is an Effective Data Generator for 3D Object Detection
- 基于高斯点云直接在3D空间放置物体,保证位置姿态准确
- 少量外部物体注入即显著提升nuScenes数据集检测性能
- 几何多样性比外观多样性更重要,硬样本无效
我们研究自动驾驶中3D目标检测的数据增强方法。利用基于高斯点云(Gaussian Splatting)的最新3D重建技术,在真实驾驶场景中直接放置3D物体,并施加显式的几何变换。与依赖鸟瞰图生成图像的扩散模型不同,本方法在重建的3D空间中直接放置物体,确保物体位置和姿态的物理合理性及精确标注。实验表明,仅引入少量外部3D物体,即可显著提升检测性能,优于现有扩散基3D增强方法。在nuScenes数据集上的测试显示,物体放置的几何多样性影响大于外观多样性。此外,通过最大化检测损失或增加视觉遮挡生成难例,对基于摄像头的3D检测增强并无明显增益。
原文摘要 · Abstract (English)
We investigate data augmentation for 3D object detection in autonomous driving. We utilize recent advancements in 3D reconstruction based on Gaussian Splatting for 3D object placement in driving scenes. Unlike existing diffusion-based methods that synthesize images conditioned on BEV layouts, our approach places 3D objects directly in the reconstructed 3D space with explicitly imposed geometric transformations. This ensures both the physical plausibility of object placement and highly accurate 3D pose and position annotations. Our experiments demonstrate that even by integrating a limited number of external 3D objects into real scenes, the augmented data significantly enhances 3D object detection performance and outperforms existing diffusion-based 3D augmentation for object detection. Extensive testing on the nuScenes dataset reveals that imposing high geometric diversity in object placement has a greater impact compared to the appearance diversity of objects. Additionally, we show that generating hard examples, either by maximizing detection loss or imposing high visual occlusion in camera images, does not lead to more efficient 3D data augmentation for camera-based 3D object detection in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。