arXiv:2505.05081cs.CV2025-05

用扩散模型实现个性化图像生成,精准保留身份特征。

PIDiff: Image Customization for Personalized Identities with Diffusion Models

  • 结合W+空间与微调策略,分离身份与背景信息
  • 生成图像保持身份特征,多样性优于基线方法
  • 适合需要精确身份控制的图像生成场景

个性化身份的文本到图像生成旨在通过文本提示和身份图像将特定身份融入图像。现有方法基于DDPM的强大生成能力,采用文本嵌入或CLIP图像嵌入表示身份信息,但未能有效解耦身份与背景信息,导致生成图像丢失关键身份特征且多样性显著下降。部分工作尝试结合StyleGAN的W+空间进行多层级特征提取,以更准确地表征身份特征,但野外图像训练中身份与背景的混合仍导致身份定位不准,产生严重的语义干扰。本文提出一种新型微调式扩散模型PIDiff,利用W+空间和定制化微调策略,避免语义纠缠,实现精准特征提取与定位。通过引入跨注意力模块与参数优化策略,PIDiff在推理时既保留身份特征,又维持预训练模型对野外图像的生成能力。实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Text-to-image generation for personalized identities aims at incorporating the specific identity into images using a text prompt and an identity image. Based on the powerful generative capabilities of DDPMs, many previous works adopt additional prompts, such as text embeddings and CLIP image embeddings, to represent the identity information, while they fail to disentangle the identity information and background information. As a result, the generated images not only lose key identity characteristics but also suffer from significantly reduced diversity. To address this issue, previous works have combined the W+ space from StyleGAN with diffusion models, leveraging this space to provide a more accurate and comprehensive representation of identity features through multi-level feature extraction. However, the entanglement of identity and background information in in-the-wild images during training prevents accurate identity localization, resulting in severe semantic interference between identity and background. In this paper, we propose a novel fine-tuning-based diffusion model for personalized identities text-to-image generation, named PIDiff, which leverages the W+ space and an identity-tailored fine-tuning strategy to avoid semantic entanglement and achieves accurate feature extraction and localization. Style editing can also be achieved by PIDiff through preserving the characteristics of identity features in the W+ space, which vary from coarse to fine. Through the combination of the proposed cross-attention block and parameter optimization strategy, PIDiff preserves the identity information and maintains the generation capability for in-the-wild images of the pre-trained model during inference. Our experimental results validate the effectiveness of our method in this task.

图像生成扩散模型身份定制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。