用视觉模型提升自动驾驶3D感知,跨传感器预训练效果显著
LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving
- 用视觉模型生成语义超像素,对齐激光雷达点云构造对比样本
- 在11个数据集上实现领先分割与检测性能,线性探测提升23.6%
- 适合做自动驾驶3D感知的模型预训练,尤其跨传感器场景
视觉基础模型(VFMs)在2D视觉感知中取得突破,但在自动驾驶所需的3D场景理解方面潜力尚未充分挖掘。本文提出LargeAD框架,支持跨多种真实驾驶数据集的大规模3D预训练。该框架利用VFMs从2D图像中提取语义丰富的超像素,并将其与激光雷达点云对齐,生成高质量的对比学习样本,从而增强2D与3D数据间的语义一致性。关键创新包括:(i) 基于VFM的超像素生成以获得精细语义表示;(ii) VFM辅助的对比学习策略实现多模态特征对齐;(iii) 超点时序一致性保障时间稳定性;(iv) 多源数据预训练提升对不同激光雷达配置的泛化能力。在11个大规模多传感器数据集上的实验表明,该方法在激光雷达分割和目标检测任务中,线性探测与微调均显著优于现有方法,展现出卓越的适应性、效率与鲁棒性。
原文摘要 · Abstract (English)
Recent advancements in vision foundation models (VFMs) have revolutionized visual perception in 2D, yet their potential for 3D scene understanding, particularly in autonomous driving applications, remains underexplored. In this paper, we introduce LargeAD, a versatile and scalable framework designed for large-scale 3D pretraining across diverse real-world driving datasets. Our framework leverages VFMs to extract semantically rich superpixels from 2D images, which are aligned with LiDAR point clouds to generate high-quality contrastive samples. This alignment facilitates cross-modal representation learning, enhancing the semantic consistency between 2D and 3D data. We introduce several key innovations: (i) VFM-driven superpixel generation for detailed semantic representation, (ii) a VFM-assisted contrastive learning strategy to align multimodal features, (iii) superpoint temporal consistency to maintain stable representations across time, and (iv) multi-source data pretraining to generalize across various LiDAR configurations. Our approach achieves substantial gains over state-of-the-art methods in linear probing and fine-tuning for LiDAR-based segmentation and object detection. Extensive experiments on 11 large-scale multi-sensor datasets highlight our superior performance, demonstrating adaptability, efficiency, and robustness in real-world autonomous driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。