arXiv:2501.04004cs.CVcs.LG2025-01CVPR被引 28

融合多种激光雷达表示,提升自动驾驶场景理解能力

LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

  • 采用专家混合框架,协同学习点云、距离图和稀疏体素等多表示
  • 在11个大规模数据集上实现更优的语义分割性能
  • 适合自动驾驶感知系统研发人员参考

激光雷达数据预训练为利用大规模现成数据集提供了可行路径,但现有方法多集中于稀疏体素表示,忽略了其他表示形式的互补优势。本文提出LiMoE框架,将专家混合(MoE)机制引入激光雷达表示学习,协同整合范围图像、稀疏体素和原始点云等多种表示。该方法包含三个阶段:i)图像到激光雷达预训练,跨表示迁移图像先验知识;ii)对比混合学习(CML),通过MoE自适应激活各表示的相关特征,并将融合特征蒸馏至统一3D网络;iii)语义混合监督(SMS),融合多表示的语义输出以提升下游分割性能。在11个大规模激光雷达数据集上的实验验证了其有效性与优越性。代码已公开。

原文摘要 · Abstract (English)

LiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlooking the complementary attributes provided by other LiDAR representations. In this work, we propose LiMoE, a framework that integrates the Mixture of Experts (MoE) paradigm into LiDAR data representation learning to synergistically combine multiple representations, such as range images, sparse voxels, and raw points. Our approach consists of three stages: i) Image-to-LiDAR Pretraining, which transfers prior knowledge from images to point clouds across different representations; ii) Contrastive Mixture Learning (CML), which uses MoE to adaptively activate relevant attributes from each representation and distills these mixed features into a unified 3D network; iii) Semantic Mixture Supervision (SMS), which combines semantic logits from multiple representations to boost downstream segmentation performance. Extensive experiments across eleven large-scale LiDAR datasets demonstrate our effectiveness and superiority. The code has been made publicly accessible.

激光雷达表示学习自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。