解决多概念图像生成中的语义纠缠问题,让角色和属性不乱搭。
SPDiffusion: Semantic Protection Diffusion Models for Multi-concept Text-to-image Generation
- 通过新方法提取概念区域,防止注意力错位
- 仅用文本提示即可实现零混淆的多概念生成
- 适合需要精准控制角色与属性的图像生成场景
近期的文本到图像模型在生成高质量图像方面取得了显著进展。然而,在多概念生成任务中,现有方法常因语义纠缠(包括概念纠缠和属性误绑定)导致文本与图像严重不符。我们发现,这种纠缠源于潜在特征某些区域错误地关注了不相关的概念或属性标记。为此,本文提出语义保护扩散模型SPDiffusion,仅依赖文本提示即可同时解决概念纠缠与属性误绑定问题。该框架引入新型概念区域提取方法SP-Extraction,以缓解跨注意力中的区域纠缠;并设计SP-Attn机制,保护概念区域免受无关属性和概念的干扰。我们在现有基准上进行评估,SPDiffusion取得当前最优性能,验证了其有效性。
原文摘要 · Abstract (English)
Recent text-to-image models have achieved impressive results in generating high-quality images. However, when tasked with multi-concept generation creating images that contain multiple characters or objects, existing methods often suffer from semantic entanglement, including concept entanglement and improper attribute binding, leading to significant text-image inconsistency. We identify that semantic entanglement arises when certain regions of the latent features attend to incorrect concept and attribute tokens. In this work, we propose the Semantic Protection Diffusion Model (SPDiffusion) to address both concept entanglement and improper attribute binding using only a text prompt as input. The SPDiffusion framework introduces a novel concept region extraction method SP-Extraction to resolve region entanglement in cross-attention, along with SP-Attn, which protects concept regions from the influence of irrelevant attributes and concepts. To evaluate our method, we test it on existing benchmarks, where SPDiffusion achieves state-of-the-art results, demonstrating its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。