让图像模型精准生成特定主体,同时忠于文字描述。
Nested Attention: Semantic-aware Attention Values for Concept Personalization

- 用嵌套注意力机制生成随查询变化的主体特征值
- 身份保留率高,且能准确响应文本提示
- 支持多主体跨领域组合生成,通用性强
将文本到图像模型个性化以在多样场景与风格中生成特定主体是快速发展的领域。现有方法常难以平衡身份保真与文本对齐。部分方法仅用单一文本标记表示主体,表达力有限;另一些则使用更丰富表示,但破坏模型先验,削弱提示对齐。本文提出嵌套注意力(Nested Attention),将丰富的图像表征注入模型现有的交叉注意力层。核心思想是通过嵌套注意力层为生成图像各区域学习选择相关主体特征,生成依赖查询的主体值。我们将该机制集成至基于编码器的个性化方法中,实验证明其可在保持高身份保真度的同时严格遵循输入文本提示。该方法具有通用性,适用于多种领域,并因保留先验能力,可实现不同领域多个个性化主体在单图中的组合生成。
原文摘要 · Abstract (English)
Personalizing text-to-image models to generate images of specific subjects across diverse scenes and styles is a rapidly advancing field. Current approaches often face challenges in maintaining a balance between identity preservation and alignment with the input text prompt. Some methods rely on a single textual token to represent a subject, which limits expressiveness, while others employ richer representations but disrupt the model's prior, diminishing prompt alignment. In this work, we introduce Nested Attention, a novel mechanism that injects a rich and expressive image representation into the model's existing cross-attention layers. Our key idea is to generate query-dependent subject values, derived from nested attention layers that learn to select relevant subject features for each region in the generated image. We integrate these nested layers into an encoder-based personalization method, and show that they enable high identity preservation while adhering to input text prompts. Our approach is general and can be trained on various domains. Additionally, its prior preservation allows us to combine multiple personalized subjects from different domains in a single image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。