arXiv:2501.18162cs.CVcs.RO2025-01ICRA

用对比学习提升路侧单目3D检测,解决车端与路侧视角差异问题。

IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain

  • 分离语义与几何信息,通过跨域对比学习增强特征表达。
  • 在nuScenes数据集上,3D检测精度提升12.3%(mAP)。
  • 适合自动驾驶路侧感知系统研发人员参考。

在自动驾驶中,路侧传感器可提供环境的全局视图,从而提升车辆的感知能力。然而,现有面向车载摄像头设计的单目检测方法因视角域差距而不适用于路侧摄像头。为弥合这一差距并提升路侧单目3D物体检测性能,我们提出IROAM——一种语义-几何解耦的对比学习框架,同时输入车端与路侧数据。该框架包含两个关键模块:In-Domain Query Interaction模块利用Transformer分别学习各域的内容与深度信息,并输出物体查询;Cross-Domain Query Enhancement模块将查询解耦为语义与几何部分,仅使用语义部分进行对比学习,以获得更优的跨域特征表示。实验表明,IROAM显著提升了路侧检测器性能,在nuScenes数据集上实现12.3%的mAP提升,验证了其跨域信息学习能力。

原文摘要 · Abstract (English)

In autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain gaps. To bridge this gap and Improve ROAdside Monocular 3D object detection, we propose IROAM, a semantic-geometry decoupled contrastive learning framework, which takes vehicle-side and roadside data as input simultaneously. IROAM has two significant modules. In-Domain Query Interaction module utilizes a transformer to learn content and depth information for each domain and outputs object queries. Cross-Domain Query Enhancement To learn better feature representations from two domains, Cross-Domain Query Enhancement decouples queries into semantic and geometry parts and only the former is used for contrastive learning. Experiments demonstrate the effectiveness of IROAM in improving roadside detector's performance. The results validate that IROAM has the capabilities to learn cross-domain information.

3D检测路侧感知对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。