arXiv:2503.06821cs.CVcs.RO2025-03中稿 · IEEE Transactions …被引 3

提出分层视角先验的通用BEV域适应框架,提升自动驾驶地图泛化能力。

HierDAMap: Towards Universal Domain Adaptive BEV Mapping via Hierarchical Perspective Priors

  • 构建全局、稀疏、实例三层视角先验,指导跨域特征对齐
  • 在Cityscapes→BDD100K上实现82.6% mIoU,优于现有方法
  • 适合做自动驾驶视觉感知与跨域地图建模的研究者

鸟瞰图(BEV)映射技术推动了自动驾驶视觉感知的创新。由于实际场景无标签,无监督域适应成为关键路径。然而现有研究对BEV映射的无监督域适应仍有限,难以覆盖所有任务。为此,本文提出HierDAMap,一种具有分层视角先验的通用全栈式BEV域适应框架。不同于仅依赖图像级先验的方法,本工作探索了全局、稀疏和实例三个层级的视角先验引导作用。框架包含三个核心组件:语义引导伪监督(SGPS)、动态感知一致性学习(DACL)和跨域视锥混合(CDFM)。SGPS通过视觉基础模型生成的2D伪标签约束跨域视角特征分布一致性;DACL利用不确定性感知的预测深度作为中介,从视角伪标签推导动态的BEV标签,以约束对应视角特征生成的粗粒度BEV特征;CDFM则利用视锥视角掩码混合双域多视角图像,通过混合后的BEV标签引导跨域视图变换与编码学习。此外,引入域内特征交换数据增强,提升域适应学习效率。代码将公开于https://github.com/lynn-yu/HierDAMap。

原文摘要 · Abstract (English)

The exploration of Bird's-Eye View (BEV) mapping technology has driven significant innovation in visual perception technology for autonomous driving. BEV mapping models need to be applied to the unlabeled real world, making the study of unsupervised domain adaptation models an essential path. However, research on unsupervised domain adaptation for BEV mapping remains limited and cannot perfectly accommodate all BEV mapping tasks. To address this gap, this paper proposes HierDAMap, a universal and holistic BEV domain adaptation framework with hierarchical perspective priors. Unlike existing research that solely focuses on image-level learning using prior knowledge, this paper explores the guiding role of perspective prior knowledge across three distinct levels: global, sparse, and instance levels. With these priors, HierDAMap consists of three essential components, including Semantic-Guided Pseudo Supervision (SGPS), Dynamic-Aware Coherence Learning (DACL), and Cross-Domain Frustum Mixing (CDFM). SGPS constrains the cross-domain consistency of perspective feature distribution through pseudo labels generated by vision foundation models in 2D space. To mitigate feature distribution discrepancies caused by spatial variations, DACL employs uncertainty-aware predicted depth as an intermediary to derive dynamic BEV labels from perspective pseudo-labels, thereby constraining the coarse BEV features derived from corresponding perspective features. CDFM, on the other hand, leverages perspective masks of the view frustum to mix multi-view perspective images from both domains, which guides cross-domain view transformation and encoding learning through mixed BEV labels. Furthermore, this paper introduces intra-domain feature exchange data augmentation to enhance the efficiency of domain adaptation learning. The source code will be made publicly available at https://github.com/lynn-yu/HierDAMap.

BEV映射域自适应自动驾驶分层先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。