用单目图像生成全景激光雷达数据,实现高精度空间控制。
Veila: Panoramic LiDAR Generation from a Monocular RGB Image
- 基于单目图像自适应融合语义与深度信息进行条件生成
- 在nuScenes等数据集上达到当前最佳生成质量与跨模态一致性
- 适合自动驾驶和机器人领域需要可控3D数据增强的研究者
真实且可控制的全景激光雷达数据生成对自动驾驶与机器人领域的可扩展3D感知至关重要。现有方法或为无条件生成(可控性差),或采用文本引导合成(缺乏精细空间控制)。利用单目RGB图像作为空间控制信号是一种低成本、可扩展的替代方案,但面临三大挑战:(i) RGB中的语义与深度线索空间分布不一,影响条件生成可靠性;(ii) RGB外观与激光雷达几何之间的模态差异,在噪声扩散过程中放大对齐误差;(iii) 单目图像与全景激光雷达在非重叠区域保持结构一致性困难。为此,我们提出Veila,一种新型条件扩散框架,包含:(1) 可靠性感知条件机制(CACM),根据局部置信度自适应平衡语义与深度信号;(2) 几何跨模态对齐(GCMA),提升噪声扩散下的鲁棒对齐能力;(3) 全景特征一致性模块(PFC),确保单目图像与全景激光雷达间全局结构一致。此外,我们引入跨模态语义一致性和深度一致性两个新指标评估对齐质量。在nuScenes、SemanticKITTI及自建KITTI-Weather基准上的实验表明,Veila在生成保真度与跨模态一致性方面均达领先水平,并能有效提升下游激光雷达语义分割性能。
原文摘要 · Abstract (English)
Realistic and controllable panoramic LiDAR data generation is critical for scalable 3D perception in autonomous driving and robotics. Existing methods either perform unconditional generation with poor controllability or adopt text-guided synthesis, which lacks fine-grained spatial control. Leveraging a monocular RGB image as a spatial control signal offers a scalable and low-cost alternative, which remains an open problem. However, it faces three core challenges: (i) semantic and depth cues from RGB are vary spatially, complicating reliable conditioning generation; (ii) modality gaps between RGB appearance and LiDAR geometry amplify alignment errors under noisy diffusion; and (iii) maintaining structural coherence between monocular RGB and panoramic LiDAR is challenging, particularly in non-overlap regions between images and LiDAR. To address these challenges, we propose Veila, a novel conditional diffusion framework that integrates: a Confidence-Aware Conditioning Mechanism (CACM) that strengthens RGB conditioning by adaptively balancing semantic and depth cues according to their local reliability; a Geometric Cross-Modal Alignment (GCMA) for robust RGB-LiDAR alignment under noisy diffusion; and a Panoramic Feature Coherence (PFC) for enforcing global structural consistency across monocular RGB and panoramic LiDAR. Additionally, we introduce two metrics, Cross-Modal Semantic Consistency and Cross-Modal Depth Consistency, to evaluate alignment quality across modalities. Experiments on nuScenes, SemanticKITTI, and our proposed KITTI-Weather benchmark demonstrate that Veila achieves state-of-the-art generation fidelity and cross-modal consistency, while enabling generative data augmentation that improves downstream LiDAR semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。