arXiv:2410.22461cs.CV2024-10NeurIPS被引 18

提出UDGA框架,让3D目标检测模型在少标注下跨域稳定工作。

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

  • 利用多视角重叠深度约束缓解视角差异带来的几何偏差。
  • 仅用1%~5%标签即可在未知目标上实现良好检测性能。
  • 适合资源受限、需快速部署的自动驾驶3D检测场景。

基于多视角相机的3D目标检测近年来在各类视觉任务中展现出实用与经济价值。然而,传统监督学习方法在面对未见且无标注的目标数据集(即直接迁移)时,常因源域与目标域间不可避免的几何错位而表现不佳。实际应用中,还面临训练资源有限及标注数据收集困难的问题。本文提出统一域泛化与自适应框架UDGA,以缓解上述挑战。首先引入多视角重叠深度约束,利用多视角间的强关联性,显著减少因视角变化导致的几何差距;其次提出标签高效域自适应方法,在仅使用1%和5%标签的情况下,仍能保持良好的源域知识,提升训练效率。整体上,UDGA框架在源域与目标域均实现稳定检测性能,有效弥合域间差异,同时大幅降低标注需求。我们在nuScenes、Lyft和Waymo等大规模基准上验证了其鲁棒性,结果优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target datasets (\ie, direct transfer) due to the inevitable geometric misalignment between the source and target domains. In practice, we also encounter constraints on resources for training models and collecting annotations for the successful deployment of 3D object detectors. In this paper, we propose Unified Domain Generalization and Adaptation (UDGA), a practical solution to mitigate those drawbacks. We first propose Multi-view Overlap Depth Constraint that leverages the strong association between multi-view, significantly alleviating geometric gaps due to perspective view changes. Then, we present a Label-Efficient Domain Adaptation approach to handle unfamiliar targets with significantly fewer amounts of labels (\ie, 1$\%$ and 5$\%)$, while preserving well-defined source knowledge for training efficiency. Overall, UDGA framework enables stable detection performance in both source and target domains, effectively bridging inevitable domain gaps, while demanding fewer annotations. We demonstrate the robustness of UDGA with large-scale benchmarks: nuScenes, Lyft, and Waymo, where our framework outperforms the current state-of-the-art methods.

3D检测域自适应多视角少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。