用图像辅助提升激光雷达语义分割的精度和稠密性
Enhancing 3D LiDAR Segmentation by Shaping Dense and Accurate 2D Semantic Predictions
- 融合相机图像,通过跨模态滤波增强2D预测稠密性
- 动态伪监督使2D预测更接近图像生成的密集语义分布
- 显著提升3D分割精度,适合城市环境感知任务
3D激光雷达点云语义分割在城市遥感中对理解真实街道环境至关重要。通过将激光雷达点云与3D语义标签投影为稀疏地图,该任务可转化为2D问题。然而,投影后的激光雷达和标签图固有的稀疏性会导致中间2D语义预测稀疏且不准确,从而限制最终3D精度。为此,我们提出一种多模态分割模型MM2D3D,利用相机图像作为辅助数据:引入跨模态引导滤波,借助相机图像提取的密集语义关系约束中间2D预测以克服标签稀疏;引入动态交叉伪监督,促使2D预测模仿相机图像生成的密集语义分布以缓解激光雷达图稀疏问题。实验表明,所提方法使中间2D预测更具稠密性和更高准确性,显著提升最终3D精度。与现有方法相比,我们在2D和3D空间均取得更优表现。
原文摘要 · Abstract (English)
Semantic segmentation of 3D LiDAR point clouds is important in urban remote sensing for understanding real-world street environments. This task, by projecting LiDAR point clouds and 3D semantic labels as sparse maps, can be reformulated as a 2D problem. However, the intrinsic sparsity of the projected LiDAR and label maps can result in sparse and inaccurate intermediate 2D semantic predictions, which in return limits the final 3D accuracy. To address this issue, we enhance this task by shaping dense and accurate 2D predictions. Specifically, we develop a multi-modal segmentation model, MM2D3D. By leveraging camera images as auxiliary data, we introduce cross-modal guided filtering to overcome label map sparsity by constraining intermediate 2D semantic predictions with dense semantic relations derived from the camera images; and we introduce dynamic cross pseudo supervision to overcome LiDAR map sparsity by encouraging the 2D predictions to emulate the dense distribution of the semantic predictions from the camera images. Experiments show that our techniques enable our model to achieve intermediate 2D semantic predictions with dense distribution and higher accuracy, which effectively enhances the final 3D accuracy. Comparisons with previous methods demonstrate our superior performance in both 2D and 3D spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。