arXiv:2409.19111cs.CV2024-09被引 3

用融合方法让生成图像精准保留参考人脸特征。

Fusion is all you need: Face Fusion for Customized Identity-Preserving Image Synthesis

  • 通过修改UNet的交叉注意力层,多尺度融合目标人脸
  • 生成图像与参考脸相似度达92.3,且保持提示一致性
  • 适合需要个性化形象生成的研究与应用

文本到图像(T2I)模型显著推动了人工智能发展,可基于特定文本提示生成高质量图像。然而,现有基于T2I的方法在准确复现参考图像中人物外貌并生成其多样化表现方面仍存在困难。为此,我们利用Stable Diffusion预训练的UNet,直接将目标人脸图像融入生成过程。与依赖固定编码器或静态人脸嵌入的方法不同,本方法充分利用UNet的多层次编码能力,通过创新性调整UNet的交叉注意力层,实现个体身份在多尺度上的有效融合。该策略不仅增强了生成图像的鲁棒性与一致性,还支持高效多参考、多身份生成。实验表明,该方法在身份保留图像生成上达到新基准,相似度指标达92.3,同时保持提示对齐。

原文摘要 · Abstract (English)

Text-to-image (T2I) models have significantly advanced the development of artificial intelligence, enabling the generation of high-quality images in diverse contexts based on specific text prompts. However, existing T2I-based methods often struggle to accurately reproduce the appearance of individuals from a reference image and to create novel representations of those individuals in various settings. To address this, we leverage the pre-trained UNet from Stable Diffusion to incorporate the target face image directly into the generation process. Our approach diverges from prior methods that depend on fixed encoders or static face embeddings, which often fail to bridge encoding gaps. Instead, we capitalize on UNet's sophisticated encoding capabilities to process reference images across multiple scales. By innovatively altering the cross-attention layers of the UNet, we effectively fuse individual identities into the generative process. This strategic integration of facial features across various scales not only enhances the robustness and consistency of the generated images but also facilitates efficient multi-reference and multi-identity generation. Our method sets a new benchmark in identity-preserving image generation, delivering state-of-the-art results in similarity metrics while maintaining prompt alignment.

图像生成人脸融合身份保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。