arXiv:2412.14821cs.CV2024-12被引 6

用极坐标与直角坐标融合提升激光雷达语义分割效率

PC-BEV: An Efficient Polar-Cartesian BEV Fusion Framework for LiDAR Semantic Segmentation

  • 在鸟瞰图空间直接融合极坐标与直角坐标分区,利用固定网格对应关系
  • 速度比传统点级方法快170倍,精度优于多视角融合方法
  • 适合需要高速实时推理的自动驾驶感知系统

尽管多视角融合在激光雷达分割中展现潜力,但其依赖计算密集的点级交互,源于范围视图与鸟瞰图间缺乏固定对应关系,限制了实际部署。本文挑战了多视角融合为高性能必要条件的普遍认知,证明通过在鸟瞰图空间内直接融合极坐标与直角坐标分区策略即可实现显著提升。所提出的纯鸟瞰图分割模型利用两种分区方式间的固有固定网格对应关系,使融合过程快达传统点级方法的170倍。同时,该方法支持稠密特征融合,相比稀疏点级方案保留更丰富的上下文信息。为兼顾场景理解与推理效率,还引入混合Transformer-CNN架构。在SemanticKITTI与nuScenes数据集上的大量实验表明,本方法在性能与推理速度上均优于此前多视角融合方法,凸显了基于鸟瞰图融合在激光雷达分割中的潜力。

原文摘要 · Abstract (English)

Although multiview fusion has demonstrated potential in LiDAR segmentation, its dependence on computationally intensive point-based interactions, arising from the lack of fixed correspondences between views such as range view and Bird's-Eye View (BEV), hinders its practical deployment. This paper challenges the prevailing notion that multiview fusion is essential for achieving high performance. We demonstrate that significant gains can be realized by directly fusing Polar and Cartesian partitioning strategies within the BEV space. Our proposed BEV-only segmentation model leverages the inherent fixed grid correspondences between these partitioning schemes, enabling a fusion process that is orders of magnitude faster (170$\times$ speedup) than conventional point-based methods. Furthermore, our approach facilitates dense feature fusion, preserving richer contextual information compared to sparse point-based alternatives. To enhance scene understanding while maintaining inference efficiency, we also introduce a hybrid Transformer-CNN architecture. Extensive evaluation on the SemanticKITTI and nuScenes datasets provides compelling evidence that our method outperforms previous multiview fusion approaches in terms of both performance and inference speed, highlighting the potential of BEV-based fusion for LiDAR segmentation. Code is available at \url{https://github.com/skyshoumeng/PC-BEV.}

激光雷达分割鸟瞰图高效融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。