arXiv:2409.02676cs.CV2024-09中稿 · the 27th IEEE Inte…被引 1

用多摄像头训练提升单摄像头的鸟瞰图感知能力

Improved Single Camera BEV Perception Using Multi-Camera Training

  • 用多视角训练时的掩码和重建损失,模拟单摄像头输入
  • 单摄像头推理下比纯单目或纯多目训练效果更好
  • 适合低成本自动驾驶系统部署,减少误判

鸟瞰图(BEV)地图预测对轨迹预测等下游自动驾驶任务至关重要。以往依赖多摄像头构建环绕视图,但大规模生产中需降低成本,使用更少摄像头成为趋势。然而,输入图像减少会导致性能下降。为此,本文提出一种方法:在训练阶段使用六摄像头数据,通过现代掩码技术、循环学习率调度和特征重建损失,使模型在单摄像头推理时仍保持高精度。该方法在单摄像头推理下优于仅用单摄像头或仅用六摄像头训练的版本,显著降低幻觉现象,提升BEV地图质量。

原文摘要 · Abstract (English)

Bird's Eye View (BEV) map prediction is essential for downstream autonomous driving tasks like trajectory prediction. In the past, this was accomplished through the use of a sophisticated sensor configuration that captured a surround view from multiple cameras. However, in large-scale production, cost efficiency is an optimization goal, so that using fewer cameras becomes more relevant. But the consequence of fewer input images correlates with a performance drop. This raises the problem of developing a BEV perception model that provides a sufficient performance on a low-cost sensor setup. Although, primarily relevant for inference time on production cars, this cost restriction is less problematic on a test vehicle during training. Therefore, the objective of our approach is to reduce the aforementioned performance drop as much as possible using a modern multi-camera surround view model reduced for single-camera inference. The approach includes three features, a modern masking technique, a cyclic Learning Rate (LR) schedule, and a feature reconstruction loss for supervising the transition from six-camera inputs to one-camera input during training. Our method outperforms versions trained strictly with one camera or strictly with six-camera surround view for single-camera inference resulting in reduced hallucination and better quality of the BEV map.

BEV感知单摄像头自动驾驶多视角训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。