arXiv:2608.06919cs.CVcs.RO2026-08

用自监督学习提升激光雷达点云表征,显著改善分割精度。

Vernata: Self-Supervised Learning of LiDAR Point Representations

论文配图:Vernata: Self-Supervised Learning of LiDAR Point Representations
图 1 · 摘自论文原文
  • 设计多教师蒸馏框架,融合2D图像语义指导点云学习。
  • 在TartanGround和Waymo上分别提升5.9和7.3点mIoU。
  • 适用于数据少或缺少颜色/法向量的场景,鲁棒性强。

激光雷达是室外机器人感知的核心模态,但深度学习模型性能受限于3D标注数据稀缺,标注成本高昂。自监督学习通过无标签数据学习通用特征缓解此问题。本文提出Vernata,一种基于Sonata架构的多模态、多教师蒸馏框架,包含三项改进:稀疏视图增强以提升对点密度变化的鲁棒性,内存库机制稳定资源受限训练,跨模态蒸馏利用高分辨率2D图像特征提供细粒度语义引导。在GrandTour、TartanGround、Waymo及自研机器人平台数据上评估,实验显示相比Sonata基线,TartanGround上mIoU达54.7(+5.9,+12.1%),Waymo上达57.1(+7.3,+14.7%)。此外,自监督方法在缺失颜色或法向量的降模态设置下仍保持优异性能,对应mIoU分别为49.4和50.2。

原文摘要 · Abstract (English)

LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain is severely limited by the scarcity of labeled data, a direct result of the high cost of 3D annotation. Self-supervised learning addresses this scarcity by learning general-purpose features from unlabeled data. In this work, we present a multi-modal, multi-teacher distillation framework for self-supervised learning on outdoor LiDAR point clouds. Building upon the Sonata architecture, we introduce Vernata, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guidance. We evaluate our method on the GrandTour, TartanGround, and Waymo datasets, as well as data collected from our own robotic platforms. Our experiments demonstrate a significant performance improvement over Sonata baselines, yielding mIoU scores of 54.7 on TartanGround (+5.9 points, +12.1%) and 57.1 on Waymo (+7.3 points, +14.7%). Finally, we show that the self-supervised approach maintains strong performance even in reduced-modality settings (lacking color or normals), achieving competitive mIoU scores of 49.4 and 50.2 on the respective datasets.

自监督学习激光雷达点云分割多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。