arXiv:2601.22376cs.CV2026-01被引 3

FlexMap无需标定即可适配任意摄像头布局,构建更可靠的高精地图。

FlexMap: Generalized HD Map Construction from Flexible Camera Configurations

  • 用跨帧注意力隐式编码三维信息,无需显式几何投影。
  • 支持多摄像头配置,缺失视角时仍保持性能稳定。
  • 适合大规模车队部署,尤其对传感器异构场景友好。

高精地图为自动驾驶系统提供道路结构的语义信息,但现有构建方法依赖标定的多摄像头系统,并需隐式或显式的2D到鸟瞰图(BEV)变换,导致在传感器故障或车辆间摄像头配置不一致时表现脆弱。本文提出FlexMap,不同于以往仅针对特定N相机阵列的方法,该方法可自适应不同摄像头配置,无需架构修改或针对每种配置重新训练。其核心创新在于摒弃显式几何投影,采用具备几何感知能力的基础模型与跨帧注意力机制,在特征空间中隐式实现三维场景理解。FlexMap包含两个关键组件:时空增强模块将跨视图空间推理与时间动态分离;相机感知解码器引入潜在相机标记,实现无需投影矩阵的视图自适应注意力。实验表明,FlexMap在多种配置下均优于现有方法,且对缺失视图和传感器差异具有鲁棒性,支持更实际的现实部署。

原文摘要 · Abstract (English)

High-definition (HD) maps provide essential semantic information of road structures for autonomous driving systems, yet current HD map construction methods require calibrated multi-camera setups and either implicit or explicit 2D-to-BEV transformations, making them fragile when sensors fail or camera configurations vary across vehicle fleets. We introduce FlexMap, unlike prior methods that are fixed to a specific N-camera rig, our approach adapts to variable camera configurations without any architectural changes or per-configuration retraining. Our key innovation eliminates explicit geometric projections by using a geometry-aware foundation model with cross-frame attention to implicitly encode 3D scene understanding in feature space. FlexMap features two core components: a spatial-temporal enhancement module that separates cross-view spatial reasoning from temporal dynamics, and a camera-aware decoder with latent camera tokens, enabling view-adaptive attention without the need for projection matrices. Experiments demonstrate that FlexMap outperforms existing methods across multiple configurations while maintaining robustness to missing views and sensor variations, enabling more practical real-world deployment.

高精地图多视角融合自适应系统自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。