arXiv:2409.14483cs.CV2024-09被引 1

通过双向互导机制,同时提升低分辨率文本识别与复原效果

One Model for Two Tasks: Cooperatively Recognizing and Recovering Low-Resolution Scene Text Images by Iterative Mutual Guidance

  • 分离训练识别与复原模型,通过迭代互导实现协同优化
  • 在两个数据集上同时超越现有方法的识别准确率与图像保真度
  • 适合需要高质量文本复原与精准识别的多任务场景

从高分辨率(HR)图像中进行场景文本识别(STR)已取得显著进展,但低分辨率(LR)图像上的文本读取仍因视觉信息不足而困难。为此,近年来提出多种场景文本图像超分辨率(STISR)模型,先生成超分辨率(SR)图像再进行识别,从而提升识别性能。然而,这些方法存在两大缺陷:一方面,STISR模型可能生成不完美甚至错误的SR图像,误导后续的STR模型;另一方面,由于STISR与STR模型联合优化,为追求高识别准确率,可能导致SR图像保真度下降。因此,识别性能与图像保真度均难以兼顾。本文提出一种新方法IMAGE(Iterative Mutual Guidance),可同步实现低分辨率场景文本图像的识别与复原。IMAGE包含专用的STR模型和定制的STISR模型,分别独立优化,并设计迭代互导机制:STR模型提供高层语义线索指导STISR模型生成更优超分图像,同时STISR模型提供底层像素线索辅助STR模型实现更准确识别。在两个低分辨率数据集上的大量实验表明,该方法在识别性能和超分辨率保真度方面均优于现有工作。

原文摘要 · Abstract (English)

Scene text recognition (STR) from high-resolution (HR) images has been significantly successful, however text reading on low-resolution (LR) images is still challenging due to insufficient visual information. Therefore, recently many scene text image super-resolution (STISR) models have been proposed to generate super-resolution (SR) images for the LR ones, then STR is done on the SR images, which thus boosts recognition performance. Nevertheless, these methods have two major weaknesses. On the one hand, STISR approaches may generate imperfect or even erroneous SR images, which mislead the subsequent recognition of STR models. On the other hand, as the STISR and STR models are jointly optimized, to pursue high recognition accuracy, the fidelity of SR images may be spoiled. As a result, neither the recognition performance nor the fidelity of STISR models are desirable. Then, can we achieve both high recognition performance and good fidelity? To this end, in this paper we propose a novel method called IMAGE (the abbreviation of Iterative MutuAl GuidancE) to effectively recognize and recover LR scene text images simultaneously. Concretely, IMAGE consists of a specialized STR model for recognition and a tailored STISR model to recover LR images, which are optimized separately. And we develop an iterative mutual guidance mechanism, with which the STR model provides high-level semantic information as clue to the STISR model for better super-resolution, meanwhile the STISR model offers essential low-level pixel clue to the STR model for more accurate recognition. Extensive experiments on two LR datasets demonstrate the superiority of our method over the existing works on both recognition performance and super-resolution fidelity.

文本识别图像复原协同优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。