arXiv:2507.18966cs.CV2025-07被引 1

用YOLO模型从真实车辆图像中提取品牌、颜色等信息,提升执法效率。

YOLO for Knowledge Extraction from Vehicle Images: A Baseline Study

  • 采用YOLO系列检测模型结合多视角推理,提升复杂场景识别准确率。
  • 在1809个车牌数据上实现最高94.86%的分类准确率,品牌识别达93.70%。
  • 小模型表现接近大模型,适合实时应用,为警务图像检索提供高效基线。

准确识别车辆的品牌、颜色和形状对执法与情报工作至关重要。本研究评估了三种前沿深度学习方法——YOLO-v11、YOLO-World和YOLO-Classification——在由新南威尔士州公路巡逻车在复杂非受控条件下采集的真实车辆图像数据集上的表现。针对品牌、形状和颜色三个任务,分别构建了超过10万张图像的数据集。通过部署多视角推理(MVI)策略提升预测性能。测试集包含29,937张图像,对应1809个车牌。不同模型规模的实验表明,最佳模型在品牌、形状、颜色及二值化颜色任务上分别达到93.70%、82.86%、85.19%和94.86%的准确率。结果表明,使用MVI可获得可用模型;且检测类模型(YOLO-v11、YOLO-World)在品牌与形状提取上优于纯分类模型。较小的YOLO变体性能与大型模型相当,显著提升实时推理效率。该研究为真实世界车辆元数据提取提供了可靠基线,可用于快速筛选和排序海量车辆图像查询。

原文摘要 · Abstract (English)

Accurate identification of vehicle attributes such as make, colour, and shape is critical for law enforcement and intelligence applications. This study evaluates the effectiveness of three state-of-the-art deep learning approaches YOLO-v11, YOLO-World, and YOLO-Classification on a real-world vehicle image dataset. This dataset was collected under challenging and unconstrained conditions by NSW Police Highway Patrol Vehicles. A multi-view inference (MVI) approach was deployed to enhance the performance of the models' predictions. To conduct the analyses, datasets with 100,000 plus images were created for each of the three metadata prediction tasks, specifically make, shape and colour. The models were tested on a separate dataset with 29,937 images belonging to 1809 number plates. Different sets of experiments have been investigated by varying the models sizes. A classification accuracy of 93.70%, 82.86%, 85.19%, and 94.86% was achieved with the best performing make, shape, colour, and colour-binary models respectively. It was concluded that there is a need to use MVI to get usable models within such complex real-world datasets. Our findings indicated that the object detection models YOLO-v11 and YOLO-World outperformed classification-only models in make and shape extraction. Moreover, smaller YOLO variants perform comparably to larger counterparts, offering substantial efficiency benefits for real-time predictions. This work provides a robust baseline for extracting vehicle metadata in real-world scenarios. Such models can be used in filtering and sorting user queries, minimising the time required to search large vehicle images datasets.

目标检测车辆识别YOLO警务应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。