解决文本到图像生成中概念耦合问题,提升个性化与文本控制的平衡。
ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion Personalization
- 通过统计分析识别概念耦合的两种依赖偏差来源
- 提出去耦损失函数,在少样本下实现更高保真度
- 适合需要精准个性化图像生成的研究者和开发者
图像个性化能够仅用少量参考图像定制文本到图像生成。然而,其核心挑战是概念耦合:有限参考图像导致模型将个性化目标与其他概念错误关联。现有方法间接应对,难以在文本控制与个性化保真度间取得理想平衡。本文通过统计分析直接揭示概念耦合源于两种不同的依赖偏差。为此,提出两个互补的即插即用损失函数:去噪去耦损失与先验去耦损失,分别最小化对应类型的依赖偏差。大量实验表明,本方法在文本控制与个性化保真度之间实现了更优权衡。
原文摘要 · Abstract (English)
Image personalization has garnered attention for its ability to customize Text-to-Image generation using only a few reference images. However, a key challenge in image personalization is the issue of conceptual coupling, where the limited number of reference images leads the model to form unwanted associations between the personalization target and other concepts. Current methods attempt to tackle this issue indirectly, leading to a suboptimal balance between text control and personalization fidelity. In this paper, we take a direct approach to the concept coupling problem through statistical analysis, revealing that it stems from two distinct sources of dependence discrepancies. We therefore propose two complementary plug-and-play loss functions: Denoising Decouple Loss and Prior Decouple loss, each designed to minimize one type of dependence discrepancy. Extensive experiments demonstrate that our approach achieves a superior trade-off between text control and personalization fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。