arXiv:2503.13743cs.CV2025-03ICRA被引 10

用一致性教师模型提升单目3D检测跨域性能,显著减少传感器差异影响。

MonoCT: Overcoming Monocular 3D Detection Domain Shift with Consistent Teacher Models

  • 引入深度增强模块与伪标签评分机制,提升伪标签精度。
  • 在六个基准上实现最高21%的AP Mod.提升,跨视角泛化能力强。
  • 适合需要跨传感器部署的自动驾驶与无人机视觉系统使用。

针对不同传感器、环境和相机设置下的单目3D目标检测问题,本文提出一种新型无监督域适应方法MonoCT,通过生成高精度伪标签实现自监督学习。受观察启发:准确的深度估计对缓解域偏移至关重要,MonoCT引入广义深度增强(GDE)模块,结合集成思想提升深度估计性能;同时提出伪标签评分(PLS)模块,利用模型内一致性测量与多样性最大化策略,进一步生成高质量伪标签用于自训练。在六个基准上的大量实验表明,MonoCT相比现有最先进方法在AP Mod.指标上最低提升21%,且在汽车、交通摄像头和无人机视角下均表现良好,具有强泛化能力。

原文摘要 · Abstract (English)

We tackle the problem of monocular 3D object detection across different sensors, environments, and camera setups. In this paper, we introduce a novel unsupervised domain adaptation approach, MonoCT, that generates highly accurate pseudo labels for self-supervision. Inspired by our observation that accurate depth estimation is critical to mitigating domain shifts, MonoCT introduces a novel Generalized Depth Enhancement (GDE) module with an ensemble concept to improve depth estimation accuracy. Moreover, we introduce a novel Pseudo Label Scoring (PLS) module by exploring inner-model consistency measurement and a Diversity Maximization (DM) strategy to further generate high-quality pseudo labels for self-training. Extensive experiments on six benchmarks show that MonoCT outperforms existing SOTA domain adaptation methods by large margins (~21% minimum for AP Mod.) and generalizes well to car, traffic camera and drone views.

单目3D检测域适应伪标签深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。