不需训练即可提升文本生成完整性,解决文本遗漏问题。
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
- 通过注意力对齐引导文本内容与图像区域匹配
- 在去噪早期应用双损失函数,显著提升召回率
- 适合需要高精度文本生成的视觉任务研究者
尽管近期进展显著,基于扩散的文生图模型在文本准确生成方面仍存在挑战。现有方法虽提出微调或无训练优化方案,但关键问题——文本遗漏(目标文本部分或完全缺失)仍未被充分关注。本文提出TextGuider,一种全新的无训练方法,通过对齐文本内容标记与图像中的文本区域,促进文本的准确且完整呈现。具体而言,我们分析了多模态扩散变换器(MM-DiT)中与文本相关的标记注意力模式,并据此在去噪初期引入两种新损失函数进行隐空间引导。该方法在测试时文本生成上达到当前最优性能,在召回率上有显著提升,同时保持了良好的OCR准确率与CLIP分数。
原文摘要 · Abstract (English)
Despite recent advances, diffusion-based text-to-image models still struggle with accurate text rendering. Several studies have proposed fine-tuning or training-free refinement methods for accurate text rendering. However, the critical issue of text omission, where the desired text is partially or entirely missing, remains largely overlooked. In this work, we propose TextGuider, a novel training-free method that encourages accurate and complete text appearance by aligning textual content tokens and text regions in the image. Specifically, we analyze attention patterns in Multi-Modal Diffusion Transformer(MM-DiT) models, particularly for text-related tokens intended to be rendered in the image. Leveraging this observation, we apply latent guidance during the early stage of denoising steps based on two loss functions that we introduce. Our method achieves state-of-the-art performance in test-time text rendering, with significant gains in recall and strong results in OCR accuracy and CLIP score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。