让文字和图片生成更一致,解决用户自定义图像时内容丢失问题
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
- 用可学习的令牌连接文本与图像先验,弥合模态差异
- 在零样本场景下超越现有方法,生成更忠实于参考图的内容
- 适合需要精准个性化图像生成的研究者与开发者
个性化图像生成旨在将用户提供的概念融入文生图模型,实现基于提示词的定制化内容生成。近期零样本方法,尤其是基于扩散Transformer的方法,通过多模态注意力机制整合参考图像信息,使生成结果同时受提示词的文本先验和参考图像的视觉先验影响。然而我们发现,当提示词与参考图像不一致时,生成结果会过度偏向文本先验,导致参考内容严重丢失。为此,我们提出AlignGen,一种跨模态先验对齐机制,通过:1)引入可学习令牌以连接文本与视觉先验;2)采用鲁棒训练策略确保先验对齐;3)在多模态注意力中使用选择性交叉模态掩码进一步对齐先验。实验表明,AlignGen优于现有零样本方法,甚至超越主流测试时优化方案。
原文摘要 · Abstract (English)
Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero-shot approaches, particularly those leveraging diffusion transformers, incorporate reference image information through multi-modal attention mechanism. This integration allows the generated output to be influenced by both the textual prior from the prompt and the visual prior from the reference image. However, we observe that when the prompt and reference image are misaligned, the generated results exhibit a stronger bias toward the textual prior, leading to a significant loss of reference content. To address this issue, we propose AlignGen, a Cross-Modality Prior Alignment mechanism that enhances personalized image generation by: 1) introducing a learnable token to bridge the gap between the textual and visual priors, 2) incorporating a robust training strategy to ensure proper prior alignment, and 3) employing a selective cross-modal attention mask within the multi-modal attention mechanism to further align the priors. Experimental results demonstrate that AlignGen outperforms existing zero-shot methods and even surpasses popular test-time optimization approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。