arXiv:2409.05834cs.CV2024-09被引 2

用2D视觉信息微调鸟瞰图模型,降低对昂贵标注数据依赖。

Vision-Driven 2D Supervised Fine-Tuning Framework for Bird's Eye View Perception

  • 基于2D语义感知结果微调鸟瞰图网络,提升泛化能力。
  • 在nuScenes和Waymo数据集上验证,性能接近传统方法。
  • 适合无激光雷达的量产自动驾驶系统部署使用。

由于出色的感知能力,视觉鸟瞰图(BEV)感知正逐步取代昂贵的激光雷达(LiDAR)感知系统,尤其在城市智能驾驶领域。然而,当前BEV感知仍需依赖LiDAR数据构建真实标签数据库,过程繁琐耗时。此外,多数量产自动驾驶系统仅配备环视摄像头,缺乏用于精确标注的LiDAR数据。为此,我们提出一种基于视觉2D语义感知的BEV感知网络微调方法,旨在提升模型在新场景数据中的泛化能力。考虑到2D感知技术的成熟度,该方法显著降低了对高成本BEV真实标签的依赖,展现出良好的工业应用前景。在nuScenes和Waymo公开数据集上的大量实验与对比分析证明了该方法的有效性。

原文摘要 · Abstract (English)

Visual bird's eye view (BEV) perception, due to its excellent perceptual capabilities, is progressively replacing costly LiDAR-based perception systems, especially in the realm of urban intelligent driving. However, this type of perception still relies on LiDAR data to construct ground truth databases, a process that is both cumbersome and time-consuming. Moreover, most massproduced autonomous driving systems are only equipped with surround camera sensors and lack LiDAR data for precise annotation. To tackle this challenge, we propose a fine-tuning method for BEV perception network based on visual 2D semantic perception, aimed at enhancing the model's generalization capabilities in new scene data. Considering the maturity and development of 2D perception technologies, our method significantly reduces the dependency on high-cost BEV ground truths and shows promising industrial application prospects. Extensive experiments and comparative analyses conducted on the nuScenes and Waymo public datasets demonstrate the effectiveness of our proposed method.

鸟瞰图感知2D到BEV自动驾驶少依赖标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。