让人脸生成更稳定,保持身份一致性和物理特征真实。
Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
- 用百万级图文掩码数据精准控制局部修改区域。
- 引入3D人脸渲染与全局特征通道,提升真实感。
- 适合需要高保真人脸编辑的图像生成场景。
在人脸图像生成中,现有文本到图像模型在响应局部语义指令时常出现面部属性泄露和物理一致性不足的问题。本文提出Face-MakeUpV2,旨在保持参考图像的人脸身份和物理特征一致性。首先,构建了包含约一百万张图像-文本-掩码对的大规模数据集FaceCaptionMask-1M,为局部语义指令提供精确的空间监督。其次,以通用文本到图像预训练模型为骨干,引入两个互补的人脸信息注入通道:3D人脸渲染通道用于融合图像的物理特征,全局人脸特征通道用于捕捉整体外观。第三,设计了两项优化目标:在模型嵌入空间中实现语义对齐以缓解属性泄露问题,并在人脸图像上使用感知损失以保持身份一致性。大量实验表明,Face-MakeUpV2在保留人脸身份和维持参考图像物理一致性方面表现最佳。结果验证了其在多样化应用中进行可靠、可控人脸编辑的实际潜力。
原文摘要 · Abstract (English)
In facial image generation, current text-to-image models often suffer from facial attribute leakage and insufficient physical consistency when responding to local semantic instructions. In this study, we propose Face-MakeUpV2, a facial image generation model that aims to maintain the consistency of face ID and physical characteristics with the reference image. First, we constructed a large-scale dataset FaceCaptionMask-1M comprising approximately one million image-text-masks pairs that provide precise spatial supervision for the local semantic instructions. Second, we employed a general text-to-image pretrained model as the backbone and introduced two complementary facial information injection channels: a 3D facial rendering channel to incorporate the physical characteristics of the image and a global facial feature channel. Third, we formulated two optimization objectives for the supervised learning of our model: semantic alignment in the model's embedding space to mitigate the attribute leakage problem and perceptual loss on facial images to preserve ID consistency. Extensive experiments demonstrated that our Face-MakeUpV2 achieves best overall performance in terms of preserving face ID and maintaining physical consistency of the reference images. These results highlight the practical potential of Face-MakeUpV2 for reliable and controllable facial editing in diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。