arXiv:2508.01387cs.CVcs.AI2025-08被引 2

用视觉语言模型实现手机拍摄车辆视频的车牌与车型识别。

Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models

  • 用多模态提示+自反思机制提升动态视频中的识别精度。
  • 在校园和公开数据集上,车牌识别准确率达91.67%、车型识别66.67%。
  • 无需专用硬件,适合移动端和移动摄像头场景应用。

自动车牌识别(ALPR)和车辆品牌型号识别是智能交通系统的核心,支持执法、收费及事故调查。将这些方法应用于手持手机或非固定车载摄像头拍摄的视频时,面临频繁相机运动、视角变化、遮挡和未知道路几何等挑战。传统ALPR依赖专用硬件和手工OCR流程,在此条件下性能下降明显。近年来大型视觉语言模型(VLMs)可直接从任意图像中识别文本和语义属性。本研究评估了VLM在手持手机和非静态车载摄像头拍摄的单目视频中进行车牌、品牌和型号识别的潜力。车牌识别管道通过筛选清晰帧,并使用多种提示策略向VLM发送多模态提示;车型识别管道采用修改后的提示并引入可选的自反思模块,该模块将查询图像与134类参考图像对比以纠正错误。在德克萨斯大学奥斯汀分校校园采集的智能手机数据集上,取得91.67%(顶1准确率)的车牌识别精度和66.67%的车型识别精度;在公开的UFPR-ALPR数据集上,分别达到83.05%和61.07%。自反思模块使车型识别平均提升5.72%。结果表明,VLM为动态交通视频分析提供了一种低成本、可扩展的解决方案。

原文摘要 · Abstract (English)

Automatic license plate recognition (ALPR) and vehicle make and model recognition underpin intelligent transportation systems, supporting law enforcement, toll collection, and post-incident investigation. Applying these methods to videos captured by handheld smartphones or non-static vehicle-mounted cameras presents unique challenges compared to fixed installations, including frequent camera motion, varying viewpoints, occlusions, and unknown road geometry. Traditional ALPR solutions, dependent on specialized hardware and handcrafted OCR pipelines, often degrade under these conditions. Recent advances in large vision-language models (VLMs) enable direct recognition of textual and semantic attributes from arbitrary imagery. This study evaluates the potential of VLMs for ALPR and makes and models recognition using monocular videos captured with handheld smartphones and non-static mounted cameras. The proposed license plate recognition pipeline filters to sharp frames, then sends a multimodal prompt to a VLM using several prompt strategies. Make and model recognition pipeline runs the same VLM with a revised prompt and an optional self-reflection module. In the self-reflection module, the model contrasts the query image with a reference from a 134-class dataset, correcting mismatches. Experiments on a smartphone dataset collected on the campus of the University of Texas at Austin, achieve top-1 accuracies of 91.67% for ALPR and 66.67% for make and model recognition. On the public UFPR-ALPR dataset, the approach attains 83.05% and 61.07%, respectively. The self-reflection module further improves results by 5.72% on average for make and model recognition. These findings demonstrate that VLMs provide a cost-effective solution for scalable, in-motion traffic video analysis.

车辆识别视觉语言模型动态视频自反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。