用多模态数据将视频分割模型迁移到4D激光雷达,实现零样本泛化分割。
Zero-Shot 4D Lidar Panoptic Segmentation
- 通过多模态传感器融合,将视频对象分割模型的伪标签迁移至4D激光雷达空间。
- 在3D零样本激光雷达全景分割上性能提升超5点PQ,首次实现4D零样本分割。
- 适合需要快速部署新物体感知的自动驾驶与机器人导航系统使用。
零样本4D激光雷达场景中任意物体的分割与识别对具身导航至关重要,应用涵盖流式感知、语义建图与定位等。然而,当前研究受限于缺乏足够多样性和标注规模的数据集。为此,我们提出SAL-4D(Lidar中的Segment Anything--4D),利用多模态机器人传感器作为桥梁,将视频对象分割(VOS)的最新进展与现成的视觉-语言基础模型结合,迁移到激光雷达领域。通过VOS模型为短时视频序列中的轨迹打伪标签,以序列级CLIP标记进行注释,并借助校准的多模态感知系统将其映射到4D激光雷达空间,实现知识蒸馏。由于预测具备时间一致性,我们在3D零样本激光雷达全景分割(LPS)上性能优于以往方法超过5 PQ,首次实现零样本4D-LPS。
原文摘要 · Abstract (English)
Zero-shot 4D segmentation and recognition of arbitrary objects in Lidar is crucial for embodied navigation, with applications ranging from streaming perception to semantic mapping and localization. However, the primary challenge in advancing research and developing generalized, versatile methods for spatio-temporal scene understanding in Lidar lies in the scarcity of datasets that provide the necessary diversity and scale of annotations.To overcome these challenges, we propose SAL-4D (Segment Anything in Lidar--4D), a method that utilizes multi-modal robotic sensor setups as a bridge to distill recent developments in Video Object Segmentation (VOS) in conjunction with off-the-shelf Vision-Language foundation models to Lidar. We utilize VOS models to pseudo-label tracklets in short video sequences, annotate these tracklets with sequence-level CLIP tokens, and lift them to the 4D Lidar space using calibrated multi-modal sensory setups to distill them to our SAL-4D model. Due to temporal consistent predictions, we outperform prior art in 3D Zero-Shot Lidar Panoptic Segmentation (LPS) over $5$ PQ, and unlock Zero-Shot 4D-LPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。