用多参考图生成高保真人脸,保持身份一致且可控
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
- 通过3D人脸关键点融合多参考图像特征
- 在FFHQ和CelebA-HQ上实现95%以上身份相似度
- 适合需要精准人脸控制的数字人生成场景
基于扩散模型的肖像生成技术在保持身份一致性方面取得进展,但使用同一身份的多张参考图时,现有方法往往生成质量较低且难以精确控制面部属性。为此,本文提出HiFi-Portrait,一种零样本高保真肖像生成方法。首先引入人脸精修模块与关键点生成器,获取细粒度的多脸特征和3D感知的关键点(包含参考身份与目标属性)。随后设计HiFi-Net,融合多脸特征并对其与关键点对齐,提升身份保真度与人脸可控性。此外,构建了基于身份的自动化数据集训练流程。大量实验表明,该方法在身份相似度与可控性上均优于当前最优方法。同时,其可兼容先前基于SDXL的工作。
原文摘要 · Abstract (English)
Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。