用对比学习提升自动驾驶鸟瞰图感知能力
BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- 在检测损失之上设计密集对比学习,优化鸟瞰图与图像特征
- nuScenes数据集上最高提升2.4%的mAP指标
- 适合关注特征表示学习的自动驾驶感知研究者
我们提出BEVCon,一种简单而有效的对比学习框架,用于提升自动驾驶中的鸟瞰图(BEV)感知性能。BEV感知提供环境的俯视表征,对3D目标检测、分割和轨迹预测至关重要。现有工作多聚焦于增强BEV编码器和任务专用头部,而我们关注尚未充分探索的表示学习潜力。BEVCon引入两个对比学习模块:实例特征对比模块用于精炼BEV特征,视角对比模块增强图像主干网络。基于检测损失设计的密集对比学习,提升了BEV编码器与主干网络的特征表示能力。在nuScenes数据集上的大量实验表明,BEVCon实现稳定性能提升,相较于最先进基线最高达+2.4% mAP。结果凸显了表示学习在BEV感知中的关键作用,并为传统任务优化提供了互补路径。
原文摘要 · Abstract (English)
We present BEVCon, a simple yet effective contrastive learning framework designed to improve Bird's Eye View (BEV) perception in autonomous driving. BEV perception offers a top-down-view representation of the surrounding environment, making it crucial for 3D object detection, segmentation, and trajectory prediction tasks. While prior work has primarily focused on enhancing BEV encoders and task-specific heads, we address the underexplored potential of representation learning in BEV models. BEVCon introduces two contrastive learning modules: an instance feature contrast module for refining BEV features and a perspective view contrast module that enhances the image backbone. The dense contrastive learning designed on top of detection losses leads to improved feature representations across both the BEV encoder and the backbone. Extensive experiments on the nuScenes dataset demonstrate that BEVCon achieves consistent performance gains, achieving up to +2.4% mAP improvement over state-of-the-art baselines. Our results highlight the critical role of representation learning in BEV perception and offer a complementary avenue to conventional task-specific optimizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。