arXiv:2505.22643cs.CV2025-05NeurIPS被引 10

Spiral用扩散模型生成带语义的激光雷达场景,效果更好且参数更少。

SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding

  • 基于范围视图的扩散模型,一步生成深度、反射率和语义图。
  • 在SemanticKITTI和nuScenes上参数最少却达到顶尖性能。
  • 生成数据可有效用于下游分割训练,大幅减少人工标注量。

借助近期的扩散模型,基于激光雷达的大规模三维场景生成已取得显著进展。尽管现有体素方法能同时生成几何结构和语义标签,但现有的范围视图方法仅能生成无标签的激光雷达场景。依赖预训练分割模型预测语义图常导致跨模态一致性不佳。为解决此问题并保留范围视图表示的优势(如计算效率高、网络设计简化),我们提出Spiral——一种新型范围视图激光雷达扩散模型,可同步生成深度图、反射率图和语义图。此外,我们引入新颖的语义感知评估指标,以衡量生成的带标签范围视图数据质量。在SemanticKITTI和nuScenes数据集上的实验表明,Spiral以最小参数量实现最先进性能,优于将生成与分割模型分步结合的两阶段方法。进一步验证显示,Spiral生成的范围图像可用于合成数据增强,在下游分割训练中显著降低激光雷达数据标注成本。

原文摘要 · Abstract (English)

Leveraging recent diffusion models, LiDAR-based large-scale 3D scene generation has achieved great success. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR scenes. Relying on pretrained segmentation models to predict the semantic maps often results in suboptimal cross-modal consistency. To address this limitation while preserving the advantages of range-view representations, such as computational efficiency and simplified network design, we propose Spiral, a novel range-view LiDAR diffusion model that simultaneously generates depth, reflectance images, and semantic maps. Furthermore, we introduce novel semantic-aware metrics to evaluate the quality of the generated labeled range-view data. Experiments on the SemanticKITTI and nuScenes datasets demonstrate that Spiral achieves state-of-the-art performance with the smallest parameter size, outperforming two-step methods that combine the generative and segmentation models. Additionally, we validate that range images generated by Spiral can be effectively used for synthetic data augmentation in the downstream segmentation training, significantly reducing the labeling effort on LiDAR data.

3D生成扩散模型激光雷达语义生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。