arXiv:2409.06827cs.CV2024-09被引 6

提出适配自动驾驶点云的跨模态对比学习方法,显著提升3D感知性能。

Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds

  • 设计实例感知与相似性平衡的对比单元,适配多模态数据差异
  • 在Waymo、nuScenes等4个数据集上实现检测与分割性能大幅提升
  • 适用于自动驾驶中多传感器融合场景,对点云模型有普适增益

车载三维环境感知依赖于激光雷达点云,但人工标注成本高。自监督预训练成为研究热点。尽管图像对比学习成功,现有方法多仅在点云单模态上训练。实际上,自动驾驶系统常集成相机与激光雷达多传感器。本文系统比较了单模态、跨模态与多模态对比学习,发现跨模态更优。针对二维图像与三维点云间巨大差异,提出专为自动驾驶点云设计的实例感知与相似性平衡对比单元。大量实验表明,该方法在四个主流基准(Waymo Open Dataset、nuScenes、SemanticKITTI、ONCE)上,对多种点云模型的3D目标检测与语义分割任务均带来显著性能提升。

原文摘要 · Abstract (English)

3D perception in LiDAR point clouds is crucial for a self-driving vehicle to properly act in 3D environment. However, manually labeling point clouds is hard and costly. There has been a growing interest in self-supervised pre-training of 3D perception models. Following the success of contrastive learning in images, current methods mostly conduct contrastive pre-training on point clouds only. Yet an autonomous driving vehicle is typically supplied with multiple sensors including cameras and LiDAR. In this context, we systematically study single modality, cross-modality, and multi-modality for contrastive learning of point clouds, and show that cross-modality wins over other alternatives. In addition, considering the huge difference between the training sources in 2D images and 3D point clouds, it remains unclear how to design more effective contrastive units for LiDAR. We therefore propose the instance-aware and similarity-balanced contrastive units that are tailored for self-driving point clouds. Extensive experiments reveal that our approach achieves remarkable performance gains over various point cloud models across the downstream perception tasks of LiDAR based 3D object detection and 3D semantic segmentation on the four popular benchmarks including Waymo Open Dataset, nuScenes, SemanticKITTI and ONCE.

点云处理自监督学习跨模态3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。