利用激光雷达时序几何一致性,无监督生成3D语义标签和检测框。
UniLiPs: Unified LiDAR Pseudo-Labeling with Geometry-Grounded Dynamic Scene Decomposition
- 基于时序累积的激光雷达地图,融合2D视觉模型语义信息生成伪标签。
- 在三个数据集上实现3D语义分割与目标检测,提升远距离深度预测51.5%。
- 无需人工标注,适合自动驾驶中海量未标注激光雷达数据的高效利用。
自动驾驶中的未标注激光雷达数据蕴含丰富的密集3D几何信息,但缺乏人工标注使其几乎无法使用,成为感知研究的主要成本障碍。本文提出一种无监督多模态伪标签方法,利用时序积累的激光雷达地图中的强几何先验,将文本与2D视觉基础模型的线索直接融合到3D空间,无需任何人工输入。方法引入一种新型迭代更新机制,强制几何与语义的一致性,同时通过不一致检测移动物体。该方法同时生成3D语义标签、3D边界框和密集激光雷达扫描,在三个数据集上表现鲁棒泛化。实验表明,相比现有伪标签方法,本方法无需额外人工监督。即使仅用少量几何一致且稠密化的激光雷达数据,也能使80-150米和150-250米范围内的深度预测误差分别降低51.5%和22.0% MAE。
原文摘要 · Abstract (English)
Unlabeled LiDAR logs, in autonomous driving applications, are inherently a gold mine of dense 3D geometry hiding in plain sight - yet they are almost useless without human labels, highlighting a dominant cost barrier for autonomous-perception research. In this work we tackle this bottleneck by leveraging temporal-geometric consistency across LiDAR sweeps to lift and fuse cues from text and 2D vision foundation models directly into 3D, without any manual input. We introduce an unsupervised multi-modal pseudo-labeling method relying on strong geometric priors learned from temporally accumulated LiDAR maps, alongside with a novel iterative update rule that enforces joint geometric-semantic consistency, and vice-versa detecting moving objects from inconsistencies. Our method simultaneously produces 3D semantic labels, 3D bounding boxes, and dense LiDAR scans, demonstrating robust generalization across three datasets. We experimentally validate that our method compares favorably to existing semantic segmentation and object detection pseudo-labeling methods, which often require additional manual supervision. We confirm that even a small fraction of our geometrically consistent, densified LiDAR improves depth prediction by 51.5% and 22.0% MAE in the 80-150 and 150-250 meters range, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。