单目路侧摄像头直接输出车辆空间位置与朝向,无需额外几何计算。
Infrastructure-based Monocular 3D Vehicle Localization Framework with Experimental Validation

- 用预训练检测器特征联合估计车辆位置、尺寸和航向角。
- 在Mcity实测中实现无后处理的车辆轨迹与朝向恢复。
- 适合城市路口基础设施感知,可扩展性强。
本文提出一种端到端学习框架,将单目路侧摄像头图像直接映射到地面固定坐标系中的车辆状态。不同于传统先检测再做几何后处理的方法,该方法利用预训练目标检测器的特征,联合估计每辆车在地平面的位置、尺寸和航向角。框架不仅用于检测,还直接进行空间与方向估计。为支持模型训练与评估,构建了基于路侧相机与无人机(UAV)同步视频的数据采集与标签生成流程。无人机作为临时俯视传感平台,提供车辆轨迹、尺寸与朝向,并转换至地面固定坐标系,与路侧相机图像时间对齐生成真值标签。在Mcity测试设施的多次实验数据上评估表明,该方法能从单目路侧图像中恢复车辆轨迹与朝向,无需独立几何后处理阶段,展现出在城市交叉口基础设施感知中可扩展的潜力。
原文摘要 · Abstract (English)
This paper presents a one-stage learning framework that maps monocular roadside-camera images directly to vehicle states in a ground-fixed coordinate frame. Unlike conventional approaches that first detect vehicles in the image plane and subsequently apply geometric post-processing, the proposed method leverages features from a pretrained object detector to jointly estimate each vehicle's ground-plane position, dimensions, and yaw angle. The framework therefore uses visual features not only for vehicle detection but also for direct spatial and orientation estimation. To support model training and evaluation, we develop a data-collection and label-generation pipeline based on synchronized video from a roadside camera and an unmanned aerial vehicle (UAV). Acting as a temporary top-view sensing platform, the UAV provides vehicle trajectories, dimensions, and orientations, which are transformed into the ground-fixed coordinate frame and temporally aligned with the roadside-camera images to generate ground-truth labels. The framework is evaluated using data collected during multiple experiments at the Mcity Test Facility. Results show that the proposed method can recover vehicle trajectories and orientations from monocular roadside imagery without a separate geometric post-processing stage, demonstrating its potential as a scalable approach to infrastructure-based perception at urban intersections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。