arXiv:2607.08896cs.CV2026-07中稿 · the IEEE ICIP 2026…

用混合注意力网络提升模糊车牌清晰度,实现极端场景下高精度识别。

HAT Super-Resolution and a PARSeq+CLIP4STR Voting Ensemble for Extreme In-the-Wild License Plate Recognition

  • 采用HAT超分辨率模型与双文本识别器集成,增强低质图像还原能力。
  • 在公开验证集上达到9.73的wECR得分,显著优于基准方案。
  • 支持实时推理,适合部署于车载或监控等严苛场景系统中。

我们提交ICIP 2026极难野外车牌超分辨率挑战赛(XLPSR)的解决方案,在公开验证榜单上取得9.73的wECR得分。系统结合混合注意力变压器(HAT)超分辨率前端、两个场景文本识别器(PARSeq-S 和 CLIP4STR-B)以及基于置信度加权的字符投票机制,对不确定位置进行弃权处理。将XLPSR视为以图像可读性为前提的识别任务:超分辨步骤旨在将字符从亚像素区域中恢复,同时显式利用不对称评分规则(+2 / -1 / 0)通过弃权策略优化性能。整个流程在RTX 3090上每序列耗时1.7秒(最坏2.7秒,第99百分位2.4秒),远低于60秒/序列的Docker资源限制。

原文摘要 · Abstract (English)

We describe our entry to the ICIP 2026 Grand Challenge on Extreme In-the-Wild License Plate Super-Resolution (XLPSR), which scored 9.73 wECR on the public validation leaderboard. The system pairs a Hybrid Attention Transformer super-resolution (HAT) front-end with an ensemble of two scene-text recognisers (PARSeq-S and CLIP4STR-B) and a confidence-weighted character-voting scheme that abstains on uncertain positions. We treat XLPSR as a recognition task gated by image legibility: the SR step exists to lift characters out of sub-pixel territory, and the asymmetric scoring rule (+2 / -1 / 0) is exploited explicitly through abstention. Our pipeline runs in 1.7 s per sequence on RTX 3090 (max 2.7 s, p99 2.4 s), well under the 60 s/sequence Docker budget.

超分辨率车牌识别视觉推理实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。