用视觉模型蒸馏训练室内激光雷达语义分割,无需人工标注。
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model

- 用视觉基础模型生成伪标签,蒸馏到激光雷达模型中。
- 伪标签下最高达56% mIoU,真实标签下约36% mIoU。
- 首次验证了跨模态蒸馏在室内场景的可行性,适合无标注数据研究者。
帧级室内激光雷达语义分割是实现高级3D场景理解与建图的关键步骤。然而,获取用于深度学习模型训练的帧级真值标签成本高昂且耗时。该问题在图像领域已通过视觉基础模型(VFMs)解决,这些模型可对图像帧进行分割。我们探索将相同方法应用于激光雷达:通过2D到3D的蒸馏流程,利用视觉基础模型处理对应的相机图像,为激光雷达扫描提供伪标签。评估基于室内SLAM数据集,采用伪标签进行下游评价;同时提供一个小规模人工标注的激光雷达数据集用于验证,因目前尚无其他带语义的室内帧级激光雷达数据集。结果表明,蒸馏模型在伪标签评估下最高达到56% mIoU,真实标签下约为36% mIoU,证明了无需人工标注即可实现跨模态蒸馏进行室内激光雷达语义分割的可行性。
原文摘要 · Abstract (English)
Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. However, acquiring frame-wise ground truth for training deep learning models is costly and time-consuming. This challenge is largely addressed, for imagery, by Visual Foundation Models (VFMs) which segment image frames. The same VFMs may be used to train a lidar scan frame segmentation model via a 2D-to-3D distillation pipeline. The success of such distillation has been shown for autonomous driving scenes, but not yet for indoor scenes. Here, we study the feasibility of repeating this success for indoor scenes, in a frame-wise distillation manner by coupling each lidar scan with a VFM-processed camera image. The evaluation is done using indoor SLAM datasets, where pseudo-labels are used for downstream evaluation. Also, a small manually annotated lidar dataset is provided for validation, as there are no other lidar frame-wise indoor datasets with semantics. Results show that the distilled model achieves up to 56% mIoU under pseudo-label evaluation and around 36% mIoU with real-label, demonstrating the feasibility of cross-modal distillation for indoor lidar semantic segmentation without manual annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。