用俯视图监督相机视角的鸟瞰表征,提升高精地图构建精度。
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

- 引入跨视角监督,将俯视图的几何拓扑先验注入摄像头的鸟瞰编码器。
- 在nuScenes上实现60×30米区域+3.9mAP,100×50米区域+9.9mAP提升。
- 无需修改推理结构或测试时输入俯视图,适合车载系统部署。
从多摄像头输入生成的鸟瞰图(BEV)表示已成为在线高精地图构建的核心接口。然而,现有方法主要依赖本车视角监督,需从不完整观测、遮挡及远距离信息稀疏中推断场景结构,受透视效应与空间稀疏性影响,难以保持一致的结构推理。本文提出跨视角监督(CVS),将本车对齐的俯视视角中的几何与拓扑先验,传递至基于摄像头的BEV编码器。CVS不添加额外语义损失,而是在共享的BEV特征空间中对齐表示,并从具有俯视优势的教师模型中蒸馏全局一致的结构知识。该监督方式增强结构一致性,且不改变推理架构,亦无需测试时输入俯视图像。在nuScenes数据集上,结合AID4AD跨视角扩展的本车对齐航拍图像进行实验,CVS在StreamMapNet基础上实现稳定提升:标准60×30米区域内+3.9mAP,扩展100×50米设置下+9.9mAP,远距离相对增益达44%。结果表明,俯视优势结构监督是提升高精地图构建中BEV表示学习的有力训练原则。
原文摘要 · Abstract (English)
Bird's-eye-view (BEV) representations derived from multi-camera input have become a central interface for online high-definition (HD) map construction. However, most approaches rely solely on ego-centric supervision, requiring large-scale scene structure to be inferred from incomplete observations, occlusions, and diminishing information density at long range, where perspective effects and spatial sparsity hinder consistent structural reasoning. We introduce Cross-View Supervision (CVS), a representation learning paradigm that transfers geometric and topological priors from an ego-aligned overhead perspective into camera-based BEV encoders. Rather than adding auxiliary semantic losses, CVS aligns representations in a shared BEV feature space and distills globally consistent structural knowledge from a perspective-privileged teacher into the ego-centric backbone. This supervision enhances structural coherence without modifying the inference architecture or requiring overhead input at test time. Experiments on nuScenes using ego-aligned aerial imagery from the AID4AD cross-view extension demonstrate consistent improvements over StreamMapNet while maintaining identical camera-only inference. CVS yields +3.9mAP in the standard $60\times30\,\mathrm{m}$ region and +9.9mAP in the extended $100\times50\,\mathrm{m}$ setting, corresponding to a 44% relative gain at long range. These results highlight perspective-privileged structural supervision as a promising training principle for improving BEV representation learning in HD map construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。