构建真实场景车牌超分辨数据集,提升低清图像识别率。
Toward Advancing License Plate Super-Resolution in Real-World Scenarios: A Dataset and Benchmark
- 用1万段真实拍摄的车牌视频构建配对高低清图像数据集。
- 超分辨结合多帧融合使识别率从1.7%提升至44.7%。
- 适合做交通监控、安防识别的模型研究者参考。
近年来,车牌识别(LPR)中的超分辨率技术致力于应对监控、交通与司法应用中低分辨率和退化图像带来的挑战。然而,现有研究依赖私有数据集和简单退化模型。为此,我们提出UFPR-SR-Plates数据集,包含10,000条车辆轨迹、100,000对真实条件下采集的低/高分辨率车牌图像。我们建立基准,每辆车提供五张连续低分辨率(LR)与五张高分辨率(HR)图像,并使用两种先进超分辨模型进行评估。同时,研究三种融合策略,分析结合领先光学字符识别(OCR)模型对多张超分辨车牌的输出如何提升整体性能。结果表明,超分辨显著提升LPR性能,采用基于字符位置多数投票(MVCP)的融合策略效果最佳:从仅用低分辨率图像的1.7%识别率,提升至超分辨后的31.1%,进一步融合五张超分辨图像输出后达44.7%。这凸显了超分辨与时序信息在真实复杂环境下提升识别精度的关键作用。该数据集已公开,可访问:https://valfride.github.io/nascimento2024toward/
原文摘要 · Abstract (English)
Recent advancements in super-resolution for License Plate Recognition (LPR) have sought to address challenges posed by low-resolution (LR) and degraded images in surveillance, traffic monitoring, and forensic applications. However, existing studies have relied on private datasets and simplistic degradation models. To address this gap, we introduce UFPR-SR-Plates, a novel dataset containing 10,000 tracks with 100,000 paired low and high-resolution license plate images captured under real-world conditions. We establish a benchmark using multiple sequential LR and high-resolution (HR) images per vehicle -- five of each -- and two state-of-the-art models for super-resolution of license plates. We also investigate three fusion strategies to evaluate how combining predictions from a leading Optical Character Recognition (OCR) model for multiple super-resolved license plates enhances overall performance. Our findings demonstrate that super-resolution significantly boosts LPR performance, with further improvements observed when applying majority vote-based fusion techniques. Specifically, the Layout-Aware and Character-Driven Network (LCDNet) model combined with the Majority Vote by Character Position (MVCP) strategy led to the highest recognition rates, increasing from 1.7% with low-resolution images to 31.1% with super-resolution, and up to 44.7% when combining OCR outputs from five super-resolved images. These findings underscore the critical role of super-resolution and temporal information in enhancing LPR accuracy under real-world, adverse conditions. The proposed dataset is publicly available to support further research and can be accessed at: https://valfride.github.io/nascimento2024toward/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。