解决多视角3D检测中跨域性能下降问题,提升自动驾驶感知鲁棒性。
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
- 设计几何感知师生框架,融合深度预测与不确定性估计增强特征提取。
- 在昼夜跨域场景下实现12.9% NDS和9.5% mAP提升,达当前最优。
- 适合关注自动驾驶视觉感知域适应的开发者与研究者。
以视觉为中心的鸟瞰图(BEV)感知在自动驾驶中具有巨大潜力。现有研究多关注效率或精度提升,却忽视了域偏移问题,导致模型迁移时性能显著下降。本文识别真实世界跨域场景中的主要域差距,首次针对多视角BEV 3D目标检测中的域适应(DA)挑战提出解决方案。由于BEV感知涉及多个组件及多几何空间(如2D、3D体素、BEV),域偏移在不同空间间累积,带来严峻挑战。为此,本文提出创新的几何感知师生框架BEVUDA++,包含可靠深度教师(RDT)与几何一致学生(GCS)模型。RDT结合目标域激光雷达与可信深度预测,基于不确定性估计生成深度感知信息,增强对体素与BEV特征的提取能力。GCS将多空间特征映射至统一几何嵌入空间,缩小两域间数据分布差异。此外,引入新型不确定性引导指数移动平均(UEMA),利用先前不确定性指导降低域偏移带来的误差累积。在四个跨域场景下进行综合实验,验证方法有效性,在昼夜适应任务中分别取得12.9% NDS与9.5% mAP提升,达到当前最佳性能。
原文摘要 · Abstract (English)
Vision-centric Bird's Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issue of domain shift has been overlooked, leading to substantial performance degradation upon transfer. We identify major domain gaps in real-world cross-domain scenarios and initiate the first effort to address the Domain Adaptation (DA) challenge in multi-view 3D object detection for BEV perception. Given the complexity of BEV perception approaches with their multiple components, domain shift accumulation across multi-geometric spaces (e.g., 2D, 3D Voxel, BEV) poses a significant challenge for BEV domain adaptation. In this paper, we introduce an innovative geometric-aware teacher-student framework, BEVUDA++, to diminish this issue, comprising a Reliable Depth Teacher (RDT) and a Geometric Consistent Student (GCS) model. Specifically, RDT effectively blends target LiDAR with dependable depth predictions to generate depth-aware information based on uncertainty estimation, enhancing the extraction of Voxel and BEV features that are essential for understanding the target domain. To collaboratively reduce the domain shift, GCS maps features from multiple spaces into a unified geometric embedding space, thereby narrowing the gap in data distribution between the two domains. Additionally, we introduce a novel Uncertainty-guided Exponential Moving Average (UEMA) to further reduce error accumulation due to domain shifts informed by previously obtained uncertainty guidance. To demonstrate the superiority of our proposed method, we execute comprehensive experiments in four cross-domain scenarios, securing state-of-the-art performance in BEV 3D object detection tasks, e.g., 12.9\% NDS and 9.5\% mAP enhancement on Day-Night adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。