无需微调,生成人脸时既保身份又准配文本描述。
Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction
- 用可控制姿态的适配器增强图像生成能力。
- 通过预测密集人脸标注提升生成质量与一致性。
- 适合需要精准人脸生成的个性化应用。
文本到图像(T2I)个性化扩散模型可根据用户输入的文本描述生成新概念图像。然而,现有方法要么需在测试时微调,要么生成结果与文本描述对齐不佳。本文提出一种新T2I个性化扩散模型Dense-Face,可在不依赖微调的情况下生成与参考主体一致身份且准确匹配文本描述的人脸图像。具体而言,我们引入一个姿态可控的适配器,以实现高保真图像生成的同时保留预训练稳定扩散(SD)的文本编辑能力。此外,利用SD UNet的内部特征预测密集人脸标注,使模型获得人脸生成领域的先验知识。实验表明,该方法在图像-文本对齐、身份保持和姿态控制方面均达到或超过当前最佳性能。
原文摘要 · Abstract (English)
The text-to-image (T2I) personalization diffusion model can generate images of the novel concept based on the user input text caption. However, existing T2I personalized methods either require test-time fine-tuning or fail to generate images that align well with the given text caption. In this work, we propose a new T2I personalization diffusion model, Dense-Face, which can generate face images with a consistent identity as the given reference subject and align well with the text caption. Specifically, we introduce a pose-controllable adapter for the high-fidelity image generation while maintaining the text-based editing ability of the pre-trained stable diffusion (SD). Additionally, we use internal features of the SD UNet to predict dense face annotations, enabling the proposed method to gain domain knowledge in face generation. Empirically, our method achieves state-of-the-art or competitive generation performance in image-text alignment, identity preservation, and pose control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。