arXiv:2410.10058cs.CV2024-10被引 7

通过丰富文本提示提升图像生成个性化泛化能力

Learning to Customize Text-to-Image Diffusion In Diverse Context

  • 仅在文本空间扩充概念上下文,不改模型结构
  • 显著提升文本-图像语义对齐,生成图像更贴合提示
  • 适配现有方法,无需额外训练成本

多数文本到图像的个性化定制方法在极少上下文的个人概念图像上微调模型,易导致过拟合,难以泛化至新上下文。现有方法依赖将个人概念有效表示为文本嵌入。本文仅通过构建丰富的文本提示集,并结合常用自监督学习目标,纯粹在文本空间多样化个人概念上下文。令人惊讶的是,该简单低成本方法显著提升了文本空间的语义对齐,此效果进一步延伸至图像空间,使生成图像对提示的忠实度更高。该方法无需架构修改,与现有文本到图像个性化方法高度兼容。我们将其与四种基线方法结合,在多个场景下均实现显著的CLIP分数提升。

原文摘要 · Abstract (English)

Most text-to-image customization techniques fine-tune models on a small set of \emph{personal concept} images captured in minimal contexts. This often results in the model becoming overfitted to these training images and unable to generalize to new contexts in future text prompts. Existing customization methods are built on the success of effectively representing personal concepts as textual embeddings. Thus, in this work, we resort to diversifying the context of these personal concepts \emph{solely} within the textual space by simply creating a contextually rich set of text prompts, together with a widely used self-supervised learning objective. Surprisingly, this straightforward and cost-effective method significantly improves semantic alignment in the textual space, and this effect further extends to the image space, resulting in higher prompt fidelity for generated images. Additionally, our approach does not require any architectural modifications, making it highly compatible with existing text-to-image customization methods. We demonstrate the broad applicability of our approach by combining it with four different baseline methods, achieving notable CLIP score improvements.

文本生成图像生成个性化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。