仅用摄像头实现高精度车载鸟瞰图感知,替代昂贵激光雷达。
Camera-Only Bird's Eye View Perception: A Neural Approach to LiDAR-Free Environmental Mapping for Autonomous Vehicles
- 融合多视角摄像头与深度估计,扩展Lift-Splat-Shoot架构生成鸟瞰图。
- 在OpenLane-V2和NuScenes上达85%道路分割准确率,车辆检测率85-90%。
- 平均定位误差仅1.2米,适合低成本自动驾驶系统部署。
自动驾驶感知系统传统依赖昂贵的激光雷达生成精确环境表示。本文提出一种纯摄像头感知框架,通过扩展Lift-Splat-Shoot架构生成鸟瞰图(BEV)。方法结合基于YOLOv11的目标检测与DepthAnythingV2单目深度估计,处理多摄像头输入,实现360度场景理解。在OpenLane-V2和NuScenes数据集上评估,道路分割准确率最高达85%,车辆检测率维持在85%-90%之间,与激光雷达真值对比时平均位置误差控制在1.2米以内。结果表明,深度学习可仅凭摄像头提取丰富空间信息,实现高性价比且不失精度的自动驾驶导航。
原文摘要 · Abstract (English)
Autonomous vehicle perception systems have traditionally relied on costly LiDAR sensors to generate precise environmental representations. In this paper, we propose a camera-only perception framework that produces Bird's Eye View (BEV) maps by extending the Lift-Splat-Shoot architecture. Our method combines YOLOv11-based object detection with DepthAnythingV2 monocular depth estimation across multi-camera inputs to achieve comprehensive 360-degree scene understanding. We evaluate our approach on the OpenLane-V2 and NuScenes datasets, achieving up to 85% road segmentation accuracy and 85-90% vehicle detection rates when compared against LiDAR ground truth, with average positional errors limited to 1.2 meters. These results highlight the potential of deep learning to extract rich spatial information using only camera inputs, enabling cost-efficient autonomous navigation without sacrificing accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。