arXiv:2410.22376cs.LGcs.AI2024-10ICLR被引 34

用大模型引导扩散模型,提升罕见概念组合生成效果

Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance

  • 利用大模型语义知识,在生成时引入相关常见概念辅助
  • 在三个数据集上最高提升28.1%的图文匹配度
  • 无需训练,适配任意预训练扩散模型和区域引导方法

当前最先进的文本到图像(T2I)扩散模型在生成罕见概念组合时表现不佳,例如具有非常规属性的物体。本文通过实证与理论分析表明,在扩散采样过程中引入与目标罕见概念相关的常见概念,可显著提升概念组合的准确性。基于此,我们提出无需训练的方法R2F,通过利用大语言模型中的丰富语义知识,规划并执行从罕见到常见的概念引导过程。该框架可兼容任意预训练扩散模型和大语言模型,并能无缝集成到区域引导扩散方法中。在包含多种罕见概念组合提示的新基准RareBench及另外两个数据集上的大量实验表明,R2F在图文对齐度上相比SD3.0和FLUX等现有模型最高提升28.1个百分点。代码已开源。

原文摘要 · Abstract (English)

State-of-the-art text-to-image (T2I) diffusion models often struggle to generate rare compositions of concepts, e.g., objects with unusual attributes. In this paper, we show that the compositional generation power of diffusion models on such rare concepts can be significantly enhanced by the Large Language Model (LLM) guidance. We start with empirical and theoretical analysis, demonstrating that exposing frequent concepts relevant to the target rare concepts during the diffusion sampling process yields more accurate concept composition. Based on this, we propose a training-free approach, R2F, that plans and executes the overall rare-to-frequent concept guidance throughout the diffusion inference by leveraging the abundant semantic knowledge in LLMs. Our framework is flexible across any pre-trained diffusion models and LLMs, and can be seamlessly integrated with the region-guided diffusion approaches. Extensive experiments on three datasets, including our newly proposed benchmark, RareBench, containing various prompts with rare compositions of concepts, R2F significantly surpasses existing models including SD3.0 and FLUX by up to 28.1%p in T2I alignment. Code is available at https://github.com/krafton-ai/Rare-to-Frequent.

扩散模型文本生成大模型引导罕见组合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。