用单目摄像头+激光雷达数据,实现铁路1公里内远距离3D物体检测。
LiDAR-Guided Monocular 3D Object Detection for Long-Range Railway Monitoring
- 基于改进YOLOv9和深度网络,融合激光雷达训练提升远距离感知。
- 在OSDaR23数据集上实现250米内物体检测,满足铁路安全需求。
- 适合铁路自动化场景,尤其适用于长距离障碍物早期预警。
德国铁路系统面临老旧基础设施挑战,亟需高自动化以安全提升列车密度。与汽车70米制动距离不同,列车需超过1公里的感知范围来提前发现轨道上的障碍物或行人。本文提出一种基于深度学习的单目图像远距离3D目标检测方法,受Faraway-Frustum启发,并在训练中引入激光雷达数据以增强深度估计。该流程包含四个模块:(1) 改进的YOLOv9用于2.5D检测,(2) 深度估计网络,(3-4) 针对短程和长程的专用3D检测头。在OSDaR23数据集上的评估表明,该方法可在250米范围内有效检测物体,展现出在铁路自动化中的应用潜力,并指出了未来优化方向。
原文摘要 · Abstract (English)
Railway systems, particularly in Germany, require high levels of automation to address legacy infrastructure challenges and increase train traffic safely. A key component of automation is robust long-range perception, essential for early hazard detection, such as obstacles at level crossings or pedestrians on tracks. Unlike automotive systems with braking distances of ~70 meters, trains require perception ranges exceeding 1 km. This paper presents an deep-learning-based approach for long-range 3D object detection tailored for autonomous trains. The method relies solely on monocular images, inspired by the Faraway-Frustum approach, and incorporates LiDAR data during training to improve depth estimation. The proposed pipeline consists of four key modules: (1) a modified YOLOv9 for 2.5D object detection, (2) a depth estimation network, and (3-4) dedicated short- and long-range 3D detection heads. Evaluations on the OSDaR23 dataset demonstrate the effectiveness of the approach in detecting objects up to 250 meters. Results highlight its potential for railway automation and outline areas for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。