用扩散模型提升真实图像超分辨率中的文字清晰度。
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders
- 引入文本感知注意力与联合分割解码器,精准恢复文字结构。
- 在多个指标上达最优,显著提升超分后文字可读性。
- 适合需要高保真文字重建的场景,如文档修复、广告牌识别。
生成模型的引入显著推进了处理真实世界退化图像的图像超分辨率(SR)技术。然而,现有方法常导致保真度问题,尤其会扭曲文字结构。本文提出一种基于扩散模型的新框架TADiSR,通过集成文本感知注意力和联合分割解码器,不仅恢复自然细节,更保障文字区域的结构准确性。同时,我们构建了完整的合成流程,生成带有精细全图文字掩码的高质量图像,融合逼真的前景文字与丰富的背景内容。大量实验表明,该方法在多个评估指标上均达到当前最优,显著提升超分后图像中文本可读性,并展现出对真实场景的强大泛化能力。代码已开源。
原文摘要 · Abstract (English)
The introduction of generative models has significantly advanced image super-resolution (SR) in handling real-world degradations. However, they often incur fidelity-related issues, particularly distorting textual structures. In this paper, we introduce a novel diffusion-based SR framework, namely TADiSR, which integrates text-aware attention and joint segmentation decoders to recover not only natural details but also the structural fidelity of text regions in degraded real-world images. Moreover, we propose a complete pipeline for synthesizing high-quality images with fine-grained full-image text masks, combining realistic foreground text regions with detailed background content. Extensive experiments demonstrate that our approach substantially enhances text legibility in super-resolved images, achieving state-of-the-art performance across multiple evaluation metrics and exhibiting strong generalization to real-world scenarios. Our code is available at \href{https://github.com/mingcv/TADiSR}{here}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。