arXiv:2510.10471cs.CVcs.LG2025-10

提升点云语义分割精度,解决伪图像表示与3D信息不一致问题

DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation

  • 通过全局-局部特征融合增强点云内部关联与上下文感知
  • 在SemanticKITTI和nuScenes上分别达到69.9%和78.7%的mIoU
  • 适合需要高精度点云分割的自动驾驶感知系统

环境感知系统对高精度地图构建与自主导航至关重要,激光雷达作为核心传感器提供精确的三维点云数据。高效处理无结构点云并提取结构化语义信息仍是重大挑战。近年来,众多基于伪图像的表示方法通过融合3D点云与2D网格,在效率与性能间取得平衡。然而,伪图像表示与原始3D信息之间的根本性不一致严重制约了2D-3D特征融合,成为信息协同融合的主要障碍,导致特征区分度差。本文提出DAGLFNet,一种基于伪图像的语义分割框架,旨在提取更具区分性的特征。其包含三个关键组件:首先,全局-局部特征融合编码(GL-FFE)模块以增强组内局部特征相关性并捕捉全局上下文信息;其次,多分支特征提取(MB-FE)网络以获取更丰富的邻域信息,提升轮廓特征的可区分性;最后,深度特征引导注意力的特征融合(FFDFA)机制以精炼跨通道特征融合精度。实验表明,DAGLFNet在SemanticKITTI和nuScenes验证集上分别取得69.9%和78.7%的平均交并比(mIoU),实现了精度与效率的优良平衡。

原文摘要 · Abstract (English)

Environmental perception systems are crucial for high-precision mapping and autonomous navigation, with LiDAR serving as a core sensor providing accurate 3D point cloud data. Efficiently processing unstructured point clouds while extracting structured semantic information remains a significant challenge. In recent years, numerous pseudo-image-based representation methods have emerged to balance efficiency and performance by fusing 3D point clouds with 2D grids. However, the fundamental inconsistency between the pseudo-image representation and the original 3D information critically undermines 2D-3D feature fusion, posing a primary obstacle for coherent information fusion and leading to poor feature discriminability. This work proposes DAGLFNet, a pseudo-image-based semantic segmentation framework designed to extract discriminative features. It incorporates three key components: first, a Global-Local Feature Fusion Encoding (GL-FFE) module to enhance intra-set local feature correlation and capture global contextual information; second, a Multi-Branch Feature Extraction (MB-FE) network to capture richer neighborhood information and improve the discriminability of contour features; and third, a Feature Fusion via Deep Feature-guided Attention (FFDFA) mechanism to refine cross-channel feature fusion precision. Experimental evaluations demonstrate that DAGLFNet achieves mean Intersection-over-Union (mIoU) scores of 69.9% and 78.7% on the validation sets of SemanticKITTI and nuScenes, respectively. The method achieves an excellent balance between accuracy and efficiency.

点云分割伪图像特征融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。