用跨视角变压器从摄像头图像生成自动驾驶鸟瞰图。
An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers

- 采用跨视角变换器学习摄像头图像到鸟瞰图的映射。
- 单城训练数据下四摄像头模型在新城镇测试中表现最稳。
- 适合自动驾驶感知系统研发人员参考。
鸟瞰图(BEV)地图为自动驾驶感知提供了结构化的俯视抽象。本文利用跨视角变换器(CVT),基于城市驾驶真实模拟器,将摄像头图像映射到三个BEV通道——道路、车道线和规划轨迹。研究评估了模型在未见过城镇的泛化能力、不同摄像头布局的影响以及两种损失函数(焦点损失与L1损失)的效果。仅使用一个城镇的数据训练时,四摄像头CVT结合L1损失在新城镇测试中表现最为稳健。整体结果凸显了CVT在将摄像头输入映射为合理准确的鸟瞰图方面的潜力。
原文摘要 · Abstract (English)
Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of different camera layouts, and two loss formulations (focal and L1). Using training data from only a town, a four-camera CVT trained with the L1 loss delivers the most robust test performance, evaluated in a new town. Overall, our results underscore CVT's promise for mapping camera inputs to reasonably accurate BEV maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。