arXiv:2506.21547cs.CVcs.RO2025-06ICCV被引 10

SAM4D实现相机与激光雷达流的可提示分割,提升自动驾驶场景下的跨模态一致性。

SAM4D: Segment Anything in Camera and LiDAR Streams

论文配图:SAM4D: Segment Anything in Camera and LiDAR Streams
图 1 · 摘自论文原文
  • 通过统一的多模态位置编码对齐相机与激光雷达特征,支持跨模态交互
  • 引入运动感知的记忆注意力机制,提升长时间序列的分割稳定性
  • 自动生成伪标签,速度远超人工标注,保留语义精度

我们提出 SAM4D,一个面向相机与激光雷达流的多模态时序基础模型,支持可提示分割。引入统一多模态位置编码(UMPE),在共享3D空间中对齐相机与激光雷达特征,实现无缝跨模态提示与交互。同时提出运动感知跨模态记忆注意力(MCMA),利用自身运动补偿增强时间一致性与长时程特征检索能力,保障动态自动驾驶场景下的鲁棒分割。为规避标注瓶颈,构建多模态自动化数据引擎,融合基于视觉-特征映射(VFM)的视频掩码片段、时空4D重建及跨模态掩码融合,以远超人工标注的速度生成相机-激光雷达对齐的伪标签,同时保持点云表征中VFM衍生的语义保真度。在自建的Waymo-4DSeg数据集上开展大量实验,验证了SAM4D强大的跨模态分割能力及其在数据标注方面的巨大潜力。

原文摘要 · Abstract (English)

We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared 3D space, enabling seamless cross-modal prompting and interaction. Additionally, we propose Motion-aware Cross-modal Memory Attention (MCMA), which leverages ego-motion compensation to enhance temporal consistency and long-horizon feature retrieval, ensuring robust segmentation across dynamically changing autonomous driving scenes. To avoid annotation bottlenecks, we develop a multi-modal automated data engine that synergizes VFM-driven video masklets, spatiotemporal 4D reconstruction, and cross-modal masklet fusion. This framework generates camera-LiDAR aligned pseudo-labels at a speed orders of magnitude faster than human annotation while preserving VFM-derived semantic fidelity in point cloud representations. We conduct extensive experiments on the constructed Waymo-4DSeg, which demonstrate the powerful cross-modal segmentation ability and great potential in data annotation of proposed SAM4D.

多模态分割自动驾驶4D感知伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。