arXiv:2603.21061cs.CV2026-03

单目摄像头实时感知系统,速度提升5.5倍且精度高

Single-Eye View: Monocular Real-time Perception Package for Autonomous Driving

  • 端到端学习结合局部建图,统一处理目标检测与深度估计
  • 在单块GPU上实现29帧/秒,比最快方法快555%
  • 适合需要低延迟的自动驾驶感知系统开发

随着基于摄像头的自动驾驶技术快速发展,性能常被优先考虑,而计算效率却常被忽视。本文提出LRHPerception,一种面向自动驾驶的单目实时感知系统,利用单视角摄像头视频理解周围环境。该系统融合端到端学习的高效性与局部建图方法的丰富表征能力,在统一框架中集成目标跟踪与预测、道路分割和深度估计。将单目图像处理为包含RGB、道路分割和像素级深度估计的五通道张量,并附加目标检测与轨迹预测结果。实验表明,系统在单块GPU上实现29 FPS的实时处理速度,相比最快基于建图的方法提速555%。

原文摘要 · Abstract (English)

Amidst the rapid advancement of camera-based autonomous driving technology, effectiveness is often prioritized with limited attention to computational efficiency. To address this issue, this paper introduces LRHPerception, a real-time monocular perception package for autonomous driving that uses single-view camera video to interpret the surrounding environment. The proposed system combines the computational efficiency of end-to-end learning with the rich representational detail of local mapping methodologies. With significant improvements in object tracking and prediction, road segmentation, and depth estimation integrated into a unified framework, LRHPerception processes monocular image data into a five-channel tensor consisting of RGB, road segmentation, and pixel-level depth estimation, augmented with object detection and trajectory prediction. Experimental results demonstrate strong performance, achieving real-time processing at 29 FPS on a single GPU, representing a 555% speedup over the fastest mapping-based approach.

自动驾驶单目感知实时系统深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。