arXiv:2503.19912cs.CVcs.LG2025-03TPAMI被引 4

提升激光雷达图像预训练中的时空一致性,增强自动驾驶感知能力。

Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data Pretraining

  • 用连续激光雷达与相机数据建模时空关系,融合运动与场景连续性。
  • 在11个异构数据集上超越当前最优方法,提升多任务泛化性能。
  • 适合研究3D感知、自动驾驶和高效预训练的开发者使用。

激光雷达表征学习已成为降低对昂贵人工标注依赖的有前景方法。现有方法主要关注激光雷达与摄像头之间的空间对齐,却常忽视驾驶场景中捕捉运动与场景连续性所必需的时间动态。为解决此问题,我们提出SuperFlow++,一种在预训练与下游任务中均整合时空线索的新框架,利用连续的激光雷达-相机数据对。SuperFlow++引入四个关键组件:(1) 视图一致性对齐模块,统一多视角相机语义信息;(2) 密度到稀疏的一致性正则化机制,提升特征在不同点云密度下的鲁棒性;(3) 基于光流的对比学习方法,建模时间关系以改善场景理解;(4) 时间投票策略,跨激光雷达扫描传播语义信息以提升预测一致性。在11个异构激光雷达数据集上的广泛评估表明,SuperFlow++在多种任务与驾驶条件下均优于当前最先进方法。此外,通过在预训练中扩展2D与3D主干网络,我们发现了涌现特性,为构建可扩展的3D基础模型提供了深层洞见。凭借强泛化能力与计算效率,SuperFlow++为自动驾驶中的数据高效激光雷达感知设立了新基准。代码已公开于https://github.com/Xiangxu-0103/SuperFlow。

原文摘要 · Abstract (English)

LiDAR representation learning has emerged as a promising approach to reducing reliance on costly and labor-intensive human annotations. While existing methods primarily focus on spatial alignment between LiDAR and camera sensors, they often overlook the temporal dynamics critical for capturing motion and scene continuity in driving scenarios. To address this limitation, we propose SuperFlow++, a novel framework that integrates spatiotemporal cues in both pretraining and downstream tasks using consecutive LiDAR-camera pairs. SuperFlow++ introduces four key components: (1) a view consistency alignment module to unify semantic information across camera views, (2) a dense-to-sparse consistency regularization mechanism to enhance feature robustness across varying point cloud densities, (3) a flow-based contrastive learning approach that models temporal relationships for improved scene understanding, and (4) a temporal voting strategy that propagates semantic information across LiDAR scans to improve prediction consistency. Extensive evaluations on 11 heterogeneous LiDAR datasets demonstrate that SuperFlow++ outperforms state-of-the-art methods across diverse tasks and driving conditions. Furthermore, by scaling both 2D and 3D backbones during pretraining, we uncover emergent properties that provide deeper insights into developing scalable 3D foundation models. With strong generalizability and computational efficiency, SuperFlow++ establishes a new benchmark for data-efficient LiDAR-based perception in autonomous driving. The code is publicly available at https://github.com/Xiangxu-0103/SuperFlow

激光雷达时空一致性预训练自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。