针对孟加拉国道路环境优化YOLO模型,提升本地车辆检测准确率。
Evaluating YOLO Architectures: Implications for Real-Time Vehicle Detection in Urban Environments of Bangladesh
- 在自建29类车辆数据集上评测6种YOLO变体,适配本地交通场景。
- YOLOv11x达63.7% [email protected],中等模型实现14-15毫秒推理速度。
- 解决稀有车型误检问题,为发展中国家自动驾驶提供基础支持。
在非孟加拉国数据集上训练的车辆检测系统难以准确识别当地独特车种,导致发展中国家自动驾驶技术存在关键空白。本研究在包含29种本地车辆(如Desi Nosimon、Leguna、Battery Rickshaw、CNG)的自建数据集上评估六种YOLO模型变体。数据集由手机拍摄的高分辨率图像(1920x1080)组成,经LabelImg手动标注为YOLO格式。结果显示,YOLOv11x表现最优,达到63.7% [email protected]、43.8% [email protected]:0.95、61.4%召回率与61.6% F1分数,但单图推理需45.8毫秒;中等规模模型(YOLOv8m、YOLOv11m)在62.5%和61.8% [email protected]下实现约14–15毫秒推理速度。研究发现,因数据不平衡,施工车辆与Desi Nosimons等稀有类别几乎无法识别;混淆矩阵显示,外观相似车辆(如Mini Trucks与Mini Covered Vans)常被误判。本研究为适应孟加拉国交通环境的鲁棒目标检测系统奠定基础,填补通用模型在发展中国家失效的空白。
原文摘要 · Abstract (English)
Vehicle detection systems trained on Non-Bangladeshi datasets struggle to accurately identify local vehicle types in Bangladesh's unique road environments, creating critical gaps in autonomous driving technology for developing regions. This study evaluates six YOLO model variants on a custom dataset featuring 29 distinct vehicle classes, including region-specific vehicles such as ``Desi Nosimon'', ``Leguna'', ``Battery Rickshaw'', and ``CNG''. The dataset comprises high-resolution images (1920x1080) captured across various Bangladeshi roads using mobile phone cameras and manually annotated using LabelImg with YOLO format bounding boxes. Performance evaluation revealed YOLOv11x as the top performer, achieving 63.7\% [email protected], 43.8\% [email protected]:0.95, 61.4\% recall, and 61.6\% F1-score, though requiring 45.8 milliseconds per image for inference. Medium variants (YOLOv8m, YOLOv11m) struck an optimal balance, delivering robust detection performance with [email protected] values of 62.5\% and 61.8\% respectively, while maintaining moderate inference times around 14-15 milliseconds. The study identified significant detection challenges for rare vehicle classes, with Construction Vehicles and Desi Nosimons showing near-zero accuracy due to dataset imbalances and insufficient training samples. Confusion matrices revealed frequent misclassifications between visually similar vehicles, particularly Mini Trucks versus Mini Covered Vans. This research provides a foundation for developing robust object detection systems specifically adapted to Bangladesh traffic conditions, addressing critical needs in autonomous vehicle technology advancement for developing regions where conventional generic-trained models fail to perform adequately.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。