arXiv:2507.07605cs.CV2025-07被引 5

用视觉语言模型提升激光雷达开放词汇分割的准确性和稳定性

LOSC: LiDAR Open-voc Segmentation Consolidator

  • 将图像语义反投影到点云,再通过时空一致性优化标签
  • 在nuScenes和SemanticKITTI上显著超越现有零样本分割方法
  • 适合关注3D感知与多模态融合的研究者

我们研究在驾驶场景中利用基于图像的视觉-语言模型(VLMs)进行激光雷达扫描的开放词汇分割。传统方法将图像语义反投影至3D点云,但生成的点标签存在噪声且稀疏。本文提出LOS C方法,通过整合时空一致性与对图像增强的鲁棒性来优化标签。基于这些改进后的标签训练3D网络,在nuScenes和SemanticKITTI数据集上,该方法在零样本开放词汇语义分割和全景分割任务中均显著超越当前最优(SOTA)水平。代码已公开于https://github.com/valeoai/LOSC。

原文摘要 · Abstract (English)

We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projected onto 3D point clouds. Yet, resulting point labels are noisy and sparse. We consolidate these labels to enforce both spatio-temporal consistency and robustness to image-level augmentations. We then train a 3D network based on these refined labels. This simple method, called LOSC, outperforms the SOTA of zero-shot open-vocabulary semantic and panoptic segmentation on both nuScenes and SemanticKITTI, with significant margins. Code is available at https://github.com/valeoai/LOSC.

激光雷达开放词汇多模态3D分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。