针对中国农村道路,用真实与合成数据混合提升自动驾驶目标检测性能。
Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

- 构建真实与合成数据混合的14类农村交通物体检测数据集。
- 1:0.5真实-合成比例下,YOLO11m模型达0.758 [email protected]最佳效果。
- 高合成数据比例引发领域偏移,长尾物体仍难识别。
当前自动驾驶目标检测模型在复杂中国农村交通场景中面临数据稀缺与泛化能力不足的问题。为此,本文提出一个专为中文农村道路设计的真实-合成混合目标检测数据集,并系统评估13种主流检测器在不同真实-合成数据比例下的表现,为农村自动驾驶中的模型选择与数据策略提供实证依据。数据集结合河南尉氏县真实图像与Unreal Engine生成的参数化虚拟场景,定义涵盖电动三轮车、低速车(LSVs)、路边摊等区域特有元素的14类物体体系。在统一训练协议下,评估了包括YOLOv5、YOLOv8、YOLO11、YOLO26系列及RT-DETR-L在内的13个模型,在全真实基线、1:0.5和1:1三种数据配置下的性能。实验表明,适度引入合成数据(1:0.5)可有效提升检测性能,其中YOLO11m在[email protected]上达到0.758最优值;而1:1比例则因领域偏移抵消数据扩展优势。多数模型能可靠识别本地车辆,但对摊位、护栏等长尾非标准物体仍存在显著感知瓶颈。本研究为农村自动驾驶感知系统的实用部署提供了关键实证支持与新见解。
原文摘要 · Abstract (English)
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest [email protected] of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。