arXiv:2501.05892cs.CV2025-01ICCV被引 1

让文字在弯曲或倾斜背景上生成更准确,且与背景融合自然。

Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation

  • 分两路生成:一路修正语义,一路注入字形结构。
  • 在斜排、曲排等复杂布局下文本准确率显著提升。
  • 无需训练,适合需快速适配新场景的视觉生成任务。

真实图像中,因设计或布局限制,倾斜或弯曲的文字(如罐子、横幅、徽章上的文字)出现频率甚至高于平面文字。尽管扩散模型已能生成高质量视觉文本,但面对倾斜或弯曲布局时,常出现文字扭曲、图文不协调的问题,根源在于训练数据不足。本文提出无需训练的新框架STGen,可精准生成复杂布局下的视觉文本,并实现图文和谐。该框架将生成过程分为两个分支:(i) 语义修正分支,利用模型生成平面文字时蕴含的准确语义信息(包括文字本身及其背景),修正复杂布局中的语义偏差;(ii) 结构注入分支,在推理阶段引入富含字形结构的字形图像潜在表示作为条件,强化文字结构。通过有效融合先验信息,构建稳固生成基础。大量实验表明,该框架在多种视觉文本布局下均实现更高准确率和更优质量。

原文摘要 · Abstract (English)

In real-world images, slanted or curved texts, especially those on cans, banners, or badges, appear as frequently, if not more so, than flat texts due to artistic design or layout constraints. While high-quality visual text generation has become available with the advanced generative capabilities of diffusion models, these models often produce distorted text and inharmonious text background when given slanted or curved text layouts due to training data limitation. In this paper, we introduce a new training-free framework, STGen, which accurately generates visual texts in challenging scenarios (\eg, slanted or curved text layouts) while harmonizing them with the text background. Our framework decomposes the visual text generation process into two branches: (i) \textbf{Semantic Rectification Branch}, which leverages the ability in generating flat but accurate visual texts of the model to guide the generation of challenging scenarios. The generated latent of flat text is abundant in accurate semantic information related both to the text itself and its background. By incorporating this, we rectify the semantic information of the texts and harmonize the integration of the text with its background in complex layouts. (ii) \textbf{Structure Injection Branch}, which reinforces the visual text structure during inference. We incorporate the latent information of the glyph image, rich in glyph structure, as a new condition to further strengthen the text structure. To enhance image harmony, we also apply an effective combination method to merge the priors, providing a solid foundation for generation. Extensive experiments across a variety of visual text layouts demonstrate that our framework achieves superior accuracy and outstanding quality.

文本生成扩散模型视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。