用图像预训练先验提升激光雷达场景生成质量
R3DPA: Leveraging 3D Representation Alignment and RGB Pretrained Priors for LiDAR Scene Generation
- 通过对齐自监督3D特征,优化生成模型中间表示
- 在KITTI-360上达到当前最优性能,显著提升生成质量
- 仅需无条件模型即可实现物体修复与场景混合
激光雷达场景合成是解决自动驾驶等机器人任务中3D数据稀缺的新兴方案。现有方法采用扩散或流匹配模型生成真实场景,但3D数据规模远低于拥有数百万样本的RGB数据集。本文提出R3DPA,首个将图像预训练先验应用于激光雷达点云生成的方法,并利用自监督3D表征实现顶尖性能。具体包括:(i) 将生成模型中间特征与自监督3D特征对齐,显著提升生成质量;(ii) 融合大规模图像预训练生成模型的知识,缓解激光雷达数据有限问题;(iii) 推理阶段实现点云控制,仅用无条件模型即可完成物体补全和场景混合。在KITTI-360基准测试中,R3DPA达到当前最优表现。代码与预训练模型已开源。
原文摘要 · Abstract (English)
LiDAR scene synthesis is an emerging solution to scarcity in 3D data for robotic tasks such as autonomous driving. Recent approaches employ diffusion or flow matching models to generate realistic scenes, but 3D data remains limited compared to RGB datasets with millions of samples. We introduce R3DPA, the first LiDAR scene generation method to unlock image-pretrained priors for LiDAR point clouds, and leverage self-supervised 3D representations for state-of-the-art results. Specifically, we (i) align intermediate features of our generative model with self-supervised 3D features, which substantially improves generation quality; (ii) transfer knowledge from large-scale image-pretrained generative models to LiDAR generation, mitigating limited LiDAR datasets; and (iii) enable point cloud control at inference for object inpainting and scene mixing with solely an unconditional model. On the KITTI-360 benchmark R3DPA achieves state of the art performance. Code and pretrained models are available at https://github.com/valeoai/R3DPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。