用图像生成的虚拟点增强稀疏点云,提升中远距离小目标分割精度。
Multimodal Point Cloud Semantic Segmentation With Virtual Point Enhancement
- 通过图像生成虚拟点,结合自适应筛选模块提升点云密度。
- 在nuScenes上引入7.7%虚拟点,实现2.89% mIoU提升。
- 适合需要高精度3D分割的自动驾驶场景应用。
基于激光雷达的三维点云识别在多个领域已证明有效。然而,点云稀疏性和密度不均给捕捉物体细节带来挑战,尤其对中远距离及小目标而言。为此,我们提出一种基于虚拟点增强(VPE)的多模态点云语义分割方法,利用图像生成的密集但含噪虚拟点来解决此问题。直接引入这些虚拟点会增加计算负担并降低性能。因此,我们设计了一种基于空间差异的自适应过滤模块,根据密度和距离选择性提取有价值的伪点,从而增强中远距离目标的密度。随后,提出一种抗噪稀疏特征编码器,包含抗噪特征提取与细粒度特征增强机制:前者利用二维图像空间降低噪声点影响,后者通过体素内邻域点聚合与下采样体素聚合强化稀疏几何特征。在SemanticKITTI和nuScenes两个大规模基准数据集上的实验验证了方法的有效性,在nuScenes上引入7.7%虚拟点的情况下,mIoU提升2.89%。
原文摘要 · Abstract (English)
LiDAR-based 3D point cloud recognition has been proven beneficial in various applications. However, the sparsity and varying density pose a significant challenge in capturing intricate details of objects, particularly for medium-range and small targets. Therefore, we propose a multi-modal point cloud semantic segmentation method based on Virtual Point Enhancement (VPE), which integrates virtual points generated from images to address these issues. These virtual points are dense but noisy, and directly incorporating them can increase computational burden and degrade performance. Therefore, we introduce a spatial difference-driven adaptive filtering module that selectively extracts valuable pseudo points from these virtual points based on density and distance, enhancing the density of medium-range targets. Subsequently, we propose a noise-robust sparse feature encoder that incorporates noise-robust feature extraction and fine-grained feature enhancement. Noise-robust feature extraction exploits the 2D image space to reduce the impact of noisy points, while fine-grained feature enhancement boosts sparse geometric features through inner-voxel neighborhood point aggregation and downsampled voxel aggregation. The results on the SemanticKITTI and nuScenes, two large-scale benchmark data sets, have validated effectiveness, significantly improving 2.89\% mIoU with the introduction of 7.7\% virtual points on nuScenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。