arXiv:2506.11839cs.CVcs.LG2025-06

用摄像头实现低成本3D目标检测,效果接近顶尖方案

Vision-based Lifting of 2D Object Detections for Automated Driving

  • 用2D CNN处理每个2D检测的点云,降低计算开销
  • KITTI基准上性能媲美顶尖图像方法,运行时间仅为三分之一
  • 覆盖各类道路使用者,适合车载系统部署

基于图像的3D目标检测是自动驾驶的必要环节,因车载摄像头成本低且已普及。当前多数先进3D检测器依赖昂贵的激光雷达获取精确深度信息。本文提出一种仅使用摄像头的流水线,将现有视觉2D检测结果提升为3D检测,作为激光雷达的低成本替代方案。与现有方法不同,本工作不仅关注车辆,还涵盖所有类型道路使用者。据我们所知,首次采用2D CNN处理每个2D检测对应的点云,极大降低计算负担。在具有挑战性的KITTI 3D目标检测基准上,该方法性能与最先进图像方法相当,但运行时间仅为后者的三分之一。

原文摘要 · Abstract (English)

Image-based 3D object detection is an inevitable part of autonomous driving because cheap onboard cameras are already available in most modern cars. Because of the accurate depth information, currently, most state-of-the-art 3D object detectors heavily rely on LiDAR data. In this paper, we propose a pipeline which lifts the results of existing vision-based 2D algorithms to 3D detections using only cameras as a cost-effective alternative to LiDAR. In contrast to existing approaches, we focus not only on cars but on all types of road users. To the best of our knowledge, we are the first using a 2D CNN to process the point cloud for each 2D detection to keep the computational effort as low as possible. Our evaluation on the challenging KITTI 3D object detection benchmark shows results comparable to state-of-the-art image-based approaches while having a runtime of only a third.

3D检测自动驾驶视觉感知轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。