arXiv:2411.16713cs.CV2024-11被引 4

用参考图引导文本生成图像,精准还原文字和标志。

Conditional Text-to-Image Generation with Reference Guidance

  • 引入参考图像作为视觉条件,提升生成精确度。
  • 在多语言文字、标志生成上优于现有方法,参数仅2855万。
  • 适合需要高精度文字/符号生成的场景,如广告设计。

文本到图像的扩散模型在根据文本指令生成视觉效果惊艳的图像方面取得了巨大成功。尽管在生成高质量图像方面取得了显著进展,但文本到图像模型在精确渲染特定主体(如文字拼写)时仍存在困难。为解决这一挑战,本文探索使用提供特定主体视觉引导的额外图像条件来指导扩散模型生成。此外,这种参考条件使模型能够以文本分词器词汇无法充分表示的方式进行条件控制,并进一步扩展了模型对新能力的泛化能力,例如生成非英文文字拼写。我们开发了多个小规模专家插件,可高效地为Stable Diffusion模型赋予接收不同参考图像的能力。每个插件均通过定制的辅助网络和损失函数训练,适用于英文场景文字生成、多语言场景文字生成以及徽标图像生成等任务。我们的专家插件在所有任务中均表现优于现有方法,每个插件仅包含2855万可训练参数。

原文摘要 · Abstract (English)

Text-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating high-fidelity visuals, text-to-image models can still struggle with precisely rendering subjects, such as text spelling. To address this challenge, this paper explores using additional conditions of an image that provides visual guidance of the particular subjects for diffusion models to generate. In addition, this reference condition empowers the model to be conditioned in ways that the vocabularies of the text tokenizer cannot adequately represent, and further extends the model's generalization to novel capabilities such as generating non-English text spellings. We develop several small-scale expert plugins that efficiently endow a Stable Diffusion model with the capability to take different references. Each plugin is trained with auxiliary networks and loss functions customized for applications such as English scene-text generation, multi-lingual scene-text generation, and logo-image generation. Our expert plugins demonstrate superior results than the existing methods on all tasks, each containing only 28.55M trainable parameters.

文生图参考引导文字生成Stable Diffusion

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。