arXiv:2601.17340cs.CV2026-01中稿 · ICASSP 2026被引 3

提出TEXTS-Diff模型,提升真实文本图像的清晰度与可读性。

TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution

  • 利用抽象概念和具体文本区域增强视觉理解与细节还原
  • 在复杂场景下实现文本修复准确率与整体画质双提升
  • 适用于需要高保真文本重建的现实应用,如文档恢复

真实世界文本图像超分辨率旨在恢复因多种退化和文本扭曲导致的视觉质量与文字可读性。然而,现有数据集中文本图像样本稀缺,导致文本区域性能不佳;且仅包含孤立文本样本的数据集限制了背景重建质量。为此,我们构建了从真实图像中采集的大规模、高质量数据集Real-Texts,覆盖多样场景,包含中英文自然文本实例。同时提出TEXTS-Aware Diffusion Model(TEXTS-Diff),在背景与文本区域均实现高质量生成。该方法通过抽象概念提升对视觉场景中文本元素的理解,通过具体文本区域增强文字细节,有效缓解文本区域常见的失真与幻觉问题,同时保持高保真场景还原。大量实验表明,该方法在多个评价指标上达到领先水平,具备优异泛化能力与复杂场景下的文本恢复精度。代码、模型与数据集将全部开源。

原文摘要 · Abstract (English)

Real-world text image super-resolution aims to restore overall visual quality and text legibility in images suffering from diverse degradations and text distortions. However, the scarcity of text image data in existing datasets results in poor performance on text regions. In addition, datasets consisting of isolated text samples limit the quality of background reconstruction. To address these limitations, we construct Real-Texts, a large-scale, high-quality dataset collected from real-world images, which covers diverse scenarios and contains natural text instances in both Chinese and English. Additionally, we propose the TEXTS-Aware Diffusion Model (TEXTS-Diff) to achieve high-quality generation in both background and textual regions. This approach leverages abstract concepts to improve the understanding of textual elements within visual scenes and concrete text regions to enhance textual details. It mitigates distortions and hallucination artifacts commonly observed in text regions, while preserving high-quality visual scene fidelity. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple evaluation metrics, exhibiting superior generalization ability and text restoration accuracy in complex scenarios. All the code, model, and dataset will be released.

图像超分文本重建扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。