arXiv:2504.06801cs.CV2025-04CVPR被引 5

让合成物体更真实地融入场景,提升单目3D检测效果

MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection

  • 根据真实场景内容学习物体合理放置分布,生成更逼真的增强数据
  • 在KITTI和NuScenes上显著提升多个检测器性能,仅用少量数据即有效
  • 适合做单目3D检测数据增强的研究者与工程师

当前单目3D检测器受限于真实世界数据集的多样性与规模。尽管数据增强有帮助,但在户外场景中生成具有场景感知的真实增强数据仍极难实现。现有合成数据方法多关注通过改进渲染提升物体外观真实性,但我们发现物体在场景中的位置、尺寸及朝向同样关键。主要障碍在于自动确定合成物体引入实际场景时的合理放置参数。为此,我们提出MonoPlace3D,一种结合3D场景内容生成真实增强的新系统。给定背景场景,MonoPlace3D学习合理的3D边界框分布,并据此采样位置进行真实物体渲染与放置。在KITTI和NuScenes两个标准数据集上的全面评估表明,MonoPlace3D能显著提升多个现有单目3D检测器的准确率,且具备高度数据效率。

原文摘要 · Abstract (English)

Current monocular 3D detectors are held back by the limited diversity and scale of real-world datasets. While data augmentation certainly helps, it's particularly difficult to generate realistic scene-aware augmented data for outdoor settings. Most current approaches to synthetic data generation focus on realistic object appearance through improved rendering techniques. However, we show that where and how objects are positioned is just as crucial for training effective 3D monocular detectors. The key obstacle lies in automatically determining realistic object placement parameters - including position, dimensions, and directional alignment when introducing synthetic objects into actual scenes. To address this, we introduce MonoPlace3D, a novel system that considers the 3D scene content to create realistic augmentations. Specifically, given a background scene, MonoPlace3D learns a distribution over plausible 3D bounding boxes. Subsequently, we render realistic objects and place them according to the locations sampled from the learned distribution. Our comprehensive evaluation on two standard datasets KITTI and NuScenes, demonstrates that MonoPlace3D significantly improves the accuracy of multiple existing monocular 3D detectors while being highly data efficient.

3D检测数据增强单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。