arXiv:2409.08091cs.CV2024-09被引 2

用新方法让零样本人物图像生成更准更稳,细节还原更好。

EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance

  • 用预训练扩散模型做主体编码器,提升身份保留能力
  • 分阶段控制文本与主体引导,避免相互干扰
  • 仅需1%数据量就达顶尖效果,适配多个主流模型

零样本个性化图像生成旨在同时遵循文本提示和主体图像生成图像,但现有方法常难以捕捉精细细节,且在两种引导间失衡。本文发现:1)主体图像编码器的选择显著影响身份保留与训练效率;2)文本与主体引导应作用于不同去噪阶段。基于此,提出EZIGen:采用固定预训练Diffusion UNet作为主体编码器,并通过分离引导主导阶段、重访部分时间步来优化主体信息传递。该方法在SD2.1-base基础上实现多个个性化生成基准的领先性能,仅使用100倍更少的训练数据。进一步迁移至SDXL后,验证其为通用的模型无关解决方案。

原文摘要 · Abstract (English)

Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of guidance. Existing methods often struggle to capture fine-grained subject details and frequently prioritize one form of guidance over the other, resulting in suboptimal subject encoding and imbalanced generation. In this study, we uncover key insights into overcoming such drawbacks, notably that 1) the choice of the subject image encoder critically influences subject identity preservation and training efficiency, and 2) the text and subject guidance should take effect at different denoising stages. Building on these insights, we introduce a new approach, EZIGen, that employs two main components: leveraging a fixed pre-trained Diffusion UNet itself as subject encoder, following a process that balances the two guidances by separating their dominance stage and revisiting certain time steps to bootstrap subject transfer quality. Through these two components, EZIGen, initially built upon SD2.1-base, achieved state-of-the-art performances on multiple personalized generation benchmarks with a unified model, while using 100 times less training data. Moreover, by further migrating our design to SDXL, EZIGen is proven to be a versatile model-agnostic solution for personalized generation. Demo Page: zichengduan.github.io/pages/EZIGen/index.html

图像生成个性化扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。