用多任务视觉语言模型提升复杂条件下车牌识别准确率
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
- 基于PaliGemma构建可开源的VehiclePaliGemma模型,专精于模糊、倾斜车牌识别
- 在马来西亚复杂场景数据集上达到87.6%准确率,每秒处理7帧图像
- 支持多车多方向车牌同时识别,适合智能交通与公安系统应用
车牌识别(LPR)依赖摄像头与计算机视觉技术自动读取车辆牌照,用于比对数据库以识别被盗车辆、无保险驾驶者及犯罪嫌疑人,显著提升警方等机构效率。传统方法依赖光学字符识别(OCR),但受噪声、模糊、天气及字符密集等因素影响,识别难度大。现有方法在扭曲图像上仍需改进。本文评估OpenAI GPT4o、Google Gemini 1.5、PaliGemma、Meta Llama 3.2、Anthropic Claude 3.5 Sonnet、LLaVA、NVIDIA VILA、moondream2等视觉语言模型(VLM)在模糊车牌上的表现,并提出开源的VehiclePaliGemma模型,针对复杂条件下的车牌进行微调。在马来西亚多变环境下采集的车牌数据集上,VehiclePaliGemma达到87.6%的识别准确率,使用A100-80GB GPU时推理速度达7帧/秒。此外,该模型展现出多任务能力,可准确识别包含多辆不同车型和颜色车辆、位置与朝向各异的车牌。
原文摘要 · Abstract (English)
License plate recognition (LPR) involves automated systems that utilize cameras and computer vision to read vehicle license plates. Such plates collected through LPR can then be compared against databases to identify stolen vehicles, uninsured drivers, crime suspects, and more. The LPR system plays a significant role in saving time for institutions such as the police force. In the past, LPR relied heavily on Optical Character Recognition (OCR), which has been widely explored to recognize characters in images. Usually, collected plate images suffer from various limitations, including noise, blurring, weather conditions, and close characters, making the recognition complex. Existing LPR methods still require significant improvement, especially for distorted images. To fill this gap, we propose utilizing visual language models (VLMs) such as OpenAI GPT4o, Google Gemini 1.5, Google PaliGemma (Pathways Language and Image model + Gemma model), Meta Llama 3.2, Anthropic Claude 3.5 Sonnet, LLaVA, NVIDIA VILA, and moondream2 to recognize such unclear plates with close characters. This paper evaluates the VLM's capability to address the aforementioned problems. Additionally, we introduce ``VehiclePaliGemma'', a fine-tuned Open-sourced PaliGemma VLM designed to recognize plates under challenging conditions. We compared our proposed VehiclePaliGemma with state-of-the-art methods and other VLMs using a dataset of Malaysian license plates collected under complex conditions. The results indicate that VehiclePaliGemma achieved superior performance with an accuracy of 87.6\%. Moreover, it is able to predict the car's plate at a speed of 7 frames per second using A100-80GB GPU. Finally, we explored the multitasking capability of VehiclePaliGemma model to accurately identify plates containing multiple cars of various models and colors, with plates positioned and oriented in different directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。