用海量无标注数据自监督训练3D感知模型,提升自动驾驶性能。
Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving
- 利用异构数据集的无标注数据自监督预训练3D模型。
- 在3D检测、分割等任务上显著提升性能,数据越多效果越好。
- 适合研究自动驾驶3D感知与自监督学习的学者和工程师。
大规模数据预训练在自然语言处理和2D视觉领域取得显著成果,激发我们探索其在自动驾驶3D感知中的潜力。本文提出利用来自异构数据集的海量无标注数据预训练3D感知模型。设计了一种自监督预训练框架,从零开始学习有效的3D表征,并引入基于提示适配器的域适应策略以降低数据集偏差。该方法在下游任务如3D目标检测、鸟瞰图分割、3D目标追踪和占用预测中均显著提升模型性能,且随着训练数据量增加,性能持续提升,表明其对自动驾驶3D感知模型具有持续优化潜力。代码将开源,以推动社区进一步研究。
原文摘要 · Abstract (English)
The significant achievements of pre-trained models leveraging large volumes of data in the field of NLP and 2D vision inspire us to explore the potential of extensive data pre-training for 3D perception in autonomous driving. Toward this goal, this paper proposes to utilize massive unlabeled data from heterogeneous datasets to pre-train 3D perception models. We introduce a self-supervised pre-training framework that learns effective 3D representations from scratch on unlabeled data, combined with a prompt adapter based domain adaptation strategy to reduce dataset bias. The approach significantly improves model performance on downstream tasks such as 3D object detection, BEV segmentation, 3D object tracking, and occupancy prediction, and shows steady performance increase as the training data volume scales up, demonstrating the potential of continually benefit 3D perception models for autonomous driving. We will release the source code to inspire further investigations in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。