arXiv:2507.20397cs.CV2025-07被引 3

用视觉语言模型融合相机与激光雷达,实现无需人工标注的3D物体自动标记。

VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving

  • 结合激光雷达几何精度与相机语义信息,提升3D标签质量。
  • 在Nuscenes上实现52.95%的开放词汇物体发现准确率,多类别检测达46.54%。
  • 支持新类别发现,适合自动驾驶数据规模化标注场景。

自动驾驶数据收集加速,但3D标签的人工标注仍因成本高、耗时长成为瓶颈。自动标注成为可扩展替代方案,但基于激光雷达的方法受限于点云稀疏、遮挡和观测不全问题,且通常缺乏语义粒度。为此,我们提出VESPA,一种融合激光雷达几何信息与相机图像语义的多模态自动标注流水线。该方法利用视觉-语言模型(VLMs)实现在点云域内的开放词汇物体识别与检测优化,支持新类别发现,并生成高质量3D伪标签,无需真实标注或高精地图。在Nuscenes数据集上,VESPA在物体发现任务中达到52.95%的平均精度,多类别检测最高达46.54%,展现出强大的可扩展3D场景理解能力。代码将在论文接受后公开。

原文摘要 · Abstract (English)

Data collection for autonomous driving is rapidly accelerating, but manual annotation, especially for 3D labels, remains a major bottleneck due to its high cost and labor intensity. Autolabeling has emerged as a scalable alternative, allowing the generation of labels for point clouds with minimal human intervention. While LiDAR-based autolabeling methods leverage geometric information, they struggle with inherent limitations of lidar data, such as sparsity, occlusions, and incomplete object observations. Furthermore, these methods typically operate in a class-agnostic manner, offering limited semantic granularity. To address these challenges, we introduce VESPA, a multimodal autolabeling pipeline that fuses the geometric precision of LiDAR with the semantic richness of camera images. Our approach leverages vision-language models (VLMs) to enable open-vocabulary object labeling and to refine detection quality directly in the point cloud domain. VESPA supports the discovery of novel categories and produces high-quality 3D pseudolabels without requiring ground-truth annotations or HD maps. On Nuscenes dataset, VESPA achieves an AP of 52.95% for object discovery and up to 46.54% for multiclass object detection, demonstrating strong performance in scalable 3D scene understanding. Code will be available upon acceptance.

3D标注多模态视觉语言模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。