arXiv:2608.25418cs.CV2026-08

用视频模型自动标注4D激光雷达数据,大幅减少人工标注工作量。

Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

论文配图:Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models
图 1 · 摘自论文原文
  • 将2D视频模型SAM2迁移至4D激光雷达,通过多视角投影生成时序一致标签。
  • 仅需每对象点选一次,即可获得整段序列的语义与全景分割结果。
  • 在SemanticKITTI上表现接近人工标注质量,适合大规模3D场景标注需求。

4D激光雷达分割进展受限于数据。在稀疏点云序列中进行时间一致的标注成本高且难以扩展,每个新任务或领域都需重新密集标注。这促使我们思考:能否完全无需人工标注,自动生成高质量的激光雷达训练数据?为此,我们提出LiDAR-SAM2,将2D视频基础模型SAM2转化为4D激光雷达领域的可扩展监督源。数据层面,通过多视角投影与时空聚合,从SAM2的视频掩码自动生成时序一致的激光雷达级标签;模型层面,设计定制化模态接口与两阶段学习目标,使SAM2的视频分割内核适配时空激光雷达结构,实现单次点击即可追踪整个序列的对象掩码。该模型在无任何人工激光雷达标注的情况下,在SemanticKITTI上生成的语义与全景标签质量接近全人工标注水平,且基于这些标签训练的模型性能逼近全真值监督。这使得LiDAR-SAM2成为显著降低三维与四维场景理解标注负担的可扩展标注工具。

原文摘要 · Abstract (English)

Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense annotation. This motivates a simple question of whether high-quality LiDAR training data can be produced automatically, without any human labeling. To this end, we introduce LiDAR-SAM2, a framework that turns a 2D video foundation model, SAM2, into a scalable source of supervision for the 4D LiDAR domain. On the data side, it automatically generates temporally coherent LiDAR-level labels from SAM2 video masks through multi-view projection and spatio-temporal aggregation. On the modeling side, a tailored modality interface and a two-stage learning objective adapt SAM2's video segmentation kernel to spatio-temporal LiDAR structure, so that a single click per object yields a consistent mask track across the sequence. Trained with no human LiDAR annotation, LiDAR-SAM2 produces semantic and panoptic labels on SemanticKITTI that approach the quality of full human annotation from only a few points, and models trained on these labels approach the performance of full ground-truth supervision. This positions LiDAR-SAM2 as a scalable labeling tool that substantially reduces the annotation burden for 3D and 4D scene understanding.

激光雷达自动标注视频模型4D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。