arXiv:2501.02523cs.CVcs.AI2025-01被引 5

用图像+文字精准生成人脸,保持原貌同时灵活调整妆容

Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation

  • 结合图像与文本提示,利用多尺度特征保留人脸身份
  • 在400万对人脸图文数据上训练,生成效果更稳定
  • 适合需要精细控制人脸外观的研究者和开发者

人脸图像具有广泛的实际应用价值。尽管当前大规模文本-图像扩散模型具备强大的生成能力,但仅靠文本提示难以生成理想的面部图像。图像提示是合理选择,但现有方法多集中于通用领域。本文旨在优化图像化妆技术,以生成目标面部图像。具体而言:(1) 基于 LAION-Face 构建了包含 400 万张高质量人脸图像-文本对的 FaceCaptionHQ-4M 数据集,用于训练 Face-MakeUp 模型;(2) 为保持参考人脸的一致性,提取/学习多尺度内容特征与姿态特征,并将其整合进扩散模型,增强对人脸身份特征的保留能力。在两个面部相关测试数据集上的验证表明,Face-MakeUp 能取得最优综合性能。代码已开源。

原文摘要 · Abstract (English)

Facial images have extensive practical applications. Although the current large-scale text-image diffusion models exhibit strong generation capabilities, it is challenging to generate the desired facial images using only text prompt. Image prompts are a logical choice. However, current methods of this type generally focus on general domain. In this paper, we aim to optimize image makeup techniques to generate the desired facial images. Specifically, (1) we built a dataset of 4 million high-quality face image-text pairs (FaceCaptionHQ-4M) based on LAION-Face to train our Face-MakeUp model; (2) to maintain consistency with the reference facial image, we extract/learn multi-scale content features and pose features for the facial image, integrating these into the diffusion model to enhance the preservation of facial identity features for diffusion models. Validation on two face-related test datasets demonstrates that our Face-MakeUp can achieve the best comprehensive performance.All codes are available at:https://github.com/ddw2AIGROUP2CQUPT/Face-MakeUp

文本生成图像人脸生成多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。