arXiv:2603.08611cs.CVcs.RO2026-03被引 2

用视觉大模型提升稀有3D物体检测,解决自动驾驶数据长尾问题。

FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

  • 融合OWLv2和Metric3Dv2的语义与深度先验,双分支检测架构。
  • 在真实驾驶数据上显著提升稀有类别(如施工人员)检测性能。
  • 适合关注自动驾驶长尾场景、多模态融合的开发者与研究者。

为应对复杂交通环境,自动驾驶车辆需识别大量涉及弱势道路使用者或交通控制设备的语义类别。然而,许多安全关键物体(如施工人员)在常规交通条件下出现频率极低,导致仅依赖驾驶数据训练样本严重不足。近期视觉基础模型在大规模数据上训练,可作为外部先验知识以增强泛化能力。本文提出FOMO-3D,首个利用视觉基础模型进行长尾3D检测的多模态3D检测器。具体地,FOMO-3D在两阶段检测框架中,结合基于激光雷达的分支与新型基于相机的分支,利用来自OWLv2和Metric3Dv2的丰富语义与深度先验,并通过注意力机制强化图像特征融合。在真实驾驶数据上的评估表明,精心设计的多模态融合策略结合视觉基础模型的丰富先验,能显著提升长尾3D检测性能。项目主页:https://waabi.ai/fomo3d/。

原文摘要 · Abstract (English)

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker) appear infrequently in nominal traffic conditions, leading to a severe shortage of training examples from driving data alone. Recent vision foundation models, which are trained on a large corpus of data, can serve as a good source of external prior knowledge to improve generalization. We propose FOMO-3D, the first multi-modal 3D detector to leverage vision foundation models for long-tailed 3D detection. Specifically, FOMO-3D exploits rich semantic and depth priors from OWLv2 and Metric3Dv2 within a two-stage detection paradigm that first generates proposals with a LiDAR-based branch and a novel camera-based branch, and refines them with attention especially to image features from OWL. Evaluations on real-world driving data show that using rich priors from vision foundation models with careful multi-modal fusion designs leads to large gains for long-tailed 3D detection. Project website is at https://waabi.ai/fomo3d/.

3D检测长尾分布视觉大模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。