arXiv:2411.18159cs.CV2024-11CVPR被引 7

自动修复文本生成图像中的错字,提升文字准确性。

Type-R: Automatically Retouching Typos for Text-to-Image Generation

  • 后处理阶段识别并修正图像中错误的文本内容
  • 结合Stable Diffusion或Flux模型,文字渲染准确率显著提升
  • 适合需要精准文字呈现的图像生成场景

尽管近期文本生成图像模型能根据详细指令生成逼真图像,但在准确呈现图像中的文字方面仍面临挑战。本文提出在后处理阶段自动修复文本错误的方法Type-R。该方法首先识别生成图像中的拼写错误,擦除错误文字,重新生成缺失文字的文本框,并对渲染后的文字进行错别字修正。大量实验表明,Type-R与最新的文本生成模型(如Stable Diffusion或Flux)结合使用时,在保持图像质量的同时实现了最高的文字渲染准确率,且在文字准确率与图像质量的平衡上优于专注于文字生成的基线方法。

原文摘要 · Abstract (English)

While recent text-to-image models can generate photorealistic images from text prompts that reflect detailed instructions, they still face significant challenges in accurately rendering words in the image. In this paper, we propose to retouch erroneous text renderings in the post-processing pipeline. Our approach, called Type-R, identifies typographical errors in the generated image, erases the erroneous text, regenerates text boxes for missing words, and finally corrects typos in the rendered words. Through extensive experiments, we show that Type-R, in combination with the latest text-to-image models such as Stable Diffusion or Flux, achieves the highest text rendering accuracy while maintaining image quality and also outperforms text-focused generation baselines in terms of balancing text accuracy and image quality.

文本生成图像修复后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。