arXiv:2504.00410cs.CV2025-04被引 1

用非类别先验提升文本图像超分辨率,更稳定且泛化更强

NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior

  • 用中间层特征替代不稳定的类别先验,提高鲁棒性
  • 在TextZoom上提升3.5%,跨数据集泛化能力提升14.8%
  • 适用于所有依赖文本先验的超分模型,适合做通用增强

场景文本图像超分辨率(STISR)旨在提升低分辨率图像的清晰度与质量。不同于以往将场景文本视为自然图像的方法,近期研究利用预训练文本识别器提取的文本先验(TP)取得显著效果。然而存在两大问题:(1)显式的类别先验(如TP)在识别错误时会负面影响超分结果,本文揭示其不稳定性,并提出使用倒数第二层特征构建非类别先验(NCAP);(2)用于生成TP的预训练识别器在低分辨率图像上表现不佳,多数方法通过联合训练识别器与超分网络来弥合域差距,但易引发先验模态过自信现象。本文提出混合硬标签与软标签以缓解该问题。在TextZoom数据集上的实验显示性能提升3.5%,跨四个文本识别数据集的泛化性能提升14.8%。所提方法可推广至所有基于文本先验的STISR网络。

原文摘要 · Abstract (English)

Scene text image super-resolution (STISR) enhances the resolution and quality of low-resolution images. Unlike previous studies that treated scene text images as natural images, recent methods using a text prior (TP), extracted from a pre-trained text recognizer, have shown strong performance. However, two major issues emerge: (1) Explicit categorical priors, like TP, can negatively impact STISR if incorrect. We reveal that these explicit priors are unstable and propose replacing them with Non-CAtegorical Prior (NCAP) using penultimate layer representations. (2) Pre-trained recognizers used to generate TP struggle with low-resolution images. To address this, most studies jointly train the recognizer with the STISR network to bridge the domain gap between low- and high-resolution images, but this can cause an overconfidence phenomenon in the prior modality. We highlight this issue and propose a method to mitigate it by mixing hard and soft labels. Experiments on the TextZoom dataset demonstrate an improvement by 3.5%, while our method significantly enhances generalization performance by 14.8\% across four text recognition datasets. Our method generalizes to all TP-guided STISR networks.

图像超分文本先验泛化增强深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。