arXiv:2504.03292cs.CV2025-04

通过融合与精修提升多概念图像生成质量

FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement

  • 分离主体与背景,重组生成多样化训练样本
  • 局部精修损失防止相似物体属性混淆
  • 适合需要精准控制多概念生成的用户

多概念文本到图像生成仍是挑战性任务。现有方法在少量样本下易过拟合,且对类别相似主体(如两只特定狗)存在属性泄露问题。本文提出融合与精修(FaR)方法,包含两个核心贡献:概念融合技术与局部精修损失函数。概念融合通过分离参考主体与背景并重新组合,生成复合图像以增加训练数据多样性,缓解小样本导致的分布狭窄问题。局部精修损失通过将每个概念的注意力图对齐至对应区域,确保扩散模型在去噪过程中能区分相似主体,避免注意力混淆。通过同时微调特定模块,FaR 在学习新概念的同时保持已有知识。实验表明,FaR 有效防止过拟合与属性泄露,保持逼真度,并优于现有先进方法。

原文摘要 · Abstract (English)

Generating multiple new concepts remains a challenging problem in the text-to-image task. Current methods often overfit when trained on a small number of samples and struggle with attribute leakage, particularly for class-similar subjects (e.g., two specific dogs). In this paper, we introduce Fuse-and-Refine (FaR), a novel approach that tackles these challenges through two key contributions: Concept Fusion technique and Localized Refinement loss function. Concept Fusion systematically augments the training data by separating reference subjects from backgrounds and recombining them into composite images to increase diversity. This augmentation technique tackles the overfitting problem by mitigating the narrow distribution of the limited training samples. In addition, Localized Refinement loss function is introduced to preserve subject representative attributes by aligning each concept's attention map to its correct region. This approach effectively prevents attribute leakage by ensuring that the diffusion model distinguishes similar subjects without mixing their attention maps during the denoising process. By fine-tuning specific modules at the same time, FaR balances the learning of new concepts with the retention of previously learned knowledge. Empirical results show that FaR not only prevents overfitting and attribute leakage while maintaining photorealism, but also outperforms other state-of-the-art methods.

文本生成图像扩散模型多概念生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。