用航拍图监督,实现动态场景下高精度鸟瞰语义地图构建。
GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes

- 以同步航拍图为监督信号,训练车载传感器生成密集鸟瞰语义图。
- 通过伪航拍重建与自训练,提升未标注数据的标注效率与准确性。
- 支持无航拍覆盖区域的伪航拍合成,适合自动驾驶环境建模。
在几何一致、以场景为中心的表示中理解道路场景对规划与地图构建至关重要。我们提出GOLD-BEV框架,仅使用时间同步的航拍图像作为训练阶段的监督信号,从车载传感器学习包含动态目标的密集鸟瞰(BEV)语义环境地图。对齐航拍的裁剪图像提供直观的目标空间,可实现密集语义标注且人工成本极低,避免了仅依赖车载视角带来的标注模糊性。关键在于严格的航拍-地面同步机制,使俯视观测能够有效监督移动交通参与者,并缓解非同步航拍源固有的时间不一致性问题。为获得可扩展的密集真值标签,我们利用领域自适应的航拍教师生成BEV伪标签,并联合训练BEV分割模型,同时可选地加入伪航拍BEV重建以增强可解释性。最后,我们通过从车载传感器合成伪航拍BEV图像,拓展了航拍覆盖范围,支持轻量级人工标注和不确定性感知的伪标签策略,适用于未标注驾驶序列。
原文摘要 · Abstract (English)
Understanding road scenes in a geometrically consistent, scene-centric representation is crucial for planning and mapping. We present GOLD-BEV, a framework that learns dense bird's-eye-view (BEV) semantic environment maps-including dynamic agents-from ego-centric sensors, using time-synchronized aerial imagery as supervision only during training. BEV-aligned aerial crops provide an intuitive target space, enabling dense semantic annotation with minimal manual effort and avoiding the ambiguity of ego-only BEV labeling. Crucially, strict aerial-ground synchronization allows overhead observations to supervise moving traffic participants and mitigates the temporal inconsistencies inherent to non-synchronized overhead sources. To obtain scalable dense targets, we generate BEV pseudo-labels using domain-adapted aerial teachers, and jointly train BEV segmentation with optional pseudo-aerial BEV reconstruction for interpretability. Finally, we extend beyond aerial coverage by learning to synthesize pseudo-aerial BEV images from ego sensors, which support lightweight human annotation and uncertainty-aware pseudo-labeling on unlabeled drives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。