用视觉语言模型实现无需布局依赖的车牌识别,提升复杂场景下的准确率。
Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- 融合视觉与语言模型,端到端识别车牌字符并迭代优化
- 在多个国际数据集上表现优于无分割方法,抗噪声和字体变形能力强
- 适合智能交通、安防监控等需要高鲁棒性的实际应用
本文提出一种面向自动车牌识别(ALPR)的模式感知框架,可在多样化的车牌布局和复杂的现实条件下稳定运行。系统由高精度检测网络与集成变压器视觉模型及迭代语言建模机制的识别模块组成。该统一识别阶段将字符识别与后置OCR优化无缝结合,通过学习特定于车牌的结构模式和格式规则,无需显式启发式修正或手动布局分类。此设计使系统联合优化视觉与语言线索,支持迭代精炼以提升在噪声、扭曲和非常规字体下的识别准确率,并在多个国际数据集(IR-LPR、UFPR-ALPR、AOLP)上实现布局无关的识别性能。实验表明,其准确率和鲁棒性显著优于近期无分割方法,证明将模式分析嵌入识别阶段可有效融合计算机视觉与语言建模,增强智能交通与监控应用中的适应能力。
原文摘要 · Abstract (English)
This work presents a pattern-aware framework for automatic license plate recognition (ALPR), designed to operate reliably across diverse plate layouts and challenging real-world conditions. The proposed system consists of a modern, high-precision detection network followed by a recognition stage that integrates a transformer-based vision model with an iterative language modelling mechanism. This unified recognition stage performs character identification and post-OCR refinement in a seamless process, learning the structural patterns and formatting rules specific to license plates without relying on explicit heuristic corrections or manual layout classification. Through this design, the system jointly optimizes visual and linguistic cues, enables iterative refinement to improve OCR accuracy under noise, distortion, and unconventional fonts, and achieves layout-independent recognition across multiple international datasets (IR-LPR, UFPR-ALPR, AOLP). Experimental results demonstrate superior accuracy and robustness compared to recent segmentation-free approaches, highlighting how embedding pattern analysis within the recognition stage bridges computer vision and language modelling for enhanced adaptability in intelligent transportation and surveillance applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。