用扩散模型实现多色头发编辑,保留人脸特征。
HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion
- 在隐空间分离控制发色与发型,提升编辑精度。
- 支持多色发型编辑,且原有人脸颜色保持不变。
- 适合需要精细发色调整的图像生成场景。
头发编辑是通过文本描述或参考图像修改发色和发型,同时保留身份、背景和衣物等无关属性的关键图像合成任务。现有方法多基于StyleGAN,但受限于其空间分布能力,难以处理多色头发编辑及面部特征保持。本文利用潜空间扩散模型(LDM)解决发型编辑问题,提出多阶段发型融合(MHB)方法,在扩散隐空间中有效分离发色与发型控制,并训练一个形变模块以对齐目标区域的发色。为增强多色发型编辑效果,还基于多色发型数据集微调了CLIP模型。实验表明,该方法在给定文本描述或参考图像时,能有效编辑多色发型并保持面部属性,优于现有方法。
原文摘要 · Abstract (English)
Hair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。