提升扩散模型在真实文本图像超分中的保真与泛化能力。
Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios

- 分阶段采样多样化数据,稳定训练并增强泛化。
- 融合预训练超分先验与跨注意力机制,提升文字保真度。
- 用置信度动态调整文本特征权重,减少错误传播。
低分辨率文本图像修复面临巨大挑战,需同时保证图像保真度和风格真实性。现有方法在复杂场景下表现不佳:传统超分模型难以保证清晰度,而扩散模型则易失真。本文提出一种新框架,显著提升扩散模型在文本图像超分任务中的泛化能力,尤其强化保真性。首先,设计渐进式数据采样策略,在不同训练阶段引入多样图像类型,稳定收敛并提升泛化;其次,采用预训练超分先验提供强空间推理能力,增强文本信息保留;再者,引入跨注意力机制更好融合文本先验;最后,利用置信度动态调节训练中文本特征的重要性,降低错误传播。在多个真实世界数据集上的实验表明,该方法不仅生成更真实的视觉外观,还显著提升文本结构准确率。
原文摘要 · Abstract (English)
Restoring low-resolution text images presents a significant challenge, as it requires maintaining both the fidelity and stylistic realism of the text in restored images. Existing text image restoration methods often fall short in hard situations, as the traditional super-resolution models cannot guarantee clarity, while diffusion-based methods fail to maintain fidelity. In this paper, we introduce a novel framework aimed at improving the generalization ability of diffusion models for text image super-resolution (SR), especially promoting fidelity. First, we propose a progressive data sampling strategy that incorporates diverse image types at different stages of training, stabilizing the convergence and improving the generalization. For the network architecture, we leverage a pre-trained SR prior to provide robust spatial reasoning capabilities, enhancing the model's ability to preserve textual information. Additionally, we employ a cross-attention mechanism to better integrate textual priors. To further reduce errors in textual priors, we utilize confidence scores to dynamically adjust the importance of textual features during training. Extensive experiments on real-world datasets demonstrate that our approach not only produces text images with more realistic visual appearances but also improves the accuracy of text structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。