arXiv:2506.09993cs.CVcs.AI2025-06被引 6

解决图像修复中文字失真问题,让修复后的文字又准又真。

Text-Aware Image Restoration with Diffusion Models

  • 用扩散模型联合训练视觉与文字识别,提升文本区域还原能力。
  • 在10万张带复杂文字的图像上训练,显著提升文字识别准确率。
  • 适合需要高精度文字恢复的应用,如文档修复、车牌识别。

图像修复旨在恢复退化的图像,但现有的基于扩散模型的修复方法在处理文本区域时往往无法忠实还原,常生成看似合理却错误的文字图案,这种现象称为文本-图像幻觉。本文提出文本感知图像修复(TAIR)新任务,要求同时恢复视觉内容和文字准确性。为此,我们构建了包含10万张高质量场景图像的SA-Text大规模基准数据集,密集标注了多样且复杂的文本实例。此外,提出多任务扩散框架TeReDiff,将扩散模型内部特征融入文本检测模块,实现双组件联合训练,从而提取丰富文本表征,并作为后续去噪步骤的提示。大量实验表明,该方法持续优于现有先进方法,在文字识别准确率上取得显著提升。

原文摘要 · Abstract (English)

Image restoration aims to recover degraded images. However, existing diffusion-based restoration methods, despite great success in natural image restoration, often struggle to faithfully reconstruct textual regions in degraded images. Those methods frequently generate plausible but incorrect text-like patterns, a phenomenon we refer to as text-image hallucination. In this paper, we introduce Text-Aware Image Restoration (TAIR), a novel restoration task that requires the simultaneous recovery of visual contents and textual fidelity. To tackle this task, we present SA-Text, a large-scale benchmark of 100K high-quality scene images densely annotated with diverse and complex text instances. Furthermore, we propose a multi-task diffusion framework, called TeReDiff, that integrates internal features from diffusion models into a text-spotting module, enabling both components to benefit from joint training. This allows for the extraction of rich text representations, which are utilized as prompts in subsequent denoising steps. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art restoration methods, achieving significant gains in text recognition accuracy. See our project page: https://cvlab-kaist.github.io/TAIR/

图像修复扩散模型文本恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。