用统一编码生成人脸,既保身份又贴文本描述。
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
- 将人脸与文本合为统一条件输入,构建联合表征空间。
- 在真实人脸生成中,身份保留和文本匹配度显著提升。
- 适合需要精准人脸生成的图像合成场景,如数字人设计。
我们提出一种新型框架,通过多模态编码策略实现身份保持的图像生成,而非将身份特征注入预训练模型。该方法将身份与文本视为统一条件输入,引入FaceCLIP——一个学习身份与文本语义联合嵌入空间的多模态编码器。给定参考人脸与文本提示,FaceCLIP生成融合身份与文本的统一表征,用于控制基础扩散模型生成身份一致且文本对齐的图像。我们还提出一种多模态对齐算法,使用损失函数将联合表征对齐至人脸、文本与图像嵌入空间。随后,将FaceCLIP集成至Stable Diffusion XL(SDXL),构建FaceCLIP-SDXL系统。相比现有方法,该系统在真实感肖像生成中表现出更强的身份保留能力与文本相关性。大量实验验证其在定量与定性上的优越性。
原文摘要 · Abstract (English)
We propose a novel framework for ID-preserving generation using a multi-modal encoding strategy rather than injecting identity features via adapters into pre-trained models. Our method treats identity and text as a unified conditioning input. To achieve this, we introduce FaceCLIP, a multi-modal encoder that learns a joint embedding space for both identity and textual semantics. Given a reference face and a text prompt, FaceCLIP produces a unified representation that encodes both identity and text, which conditions a base diffusion model to generate images that are identity-consistent and text-aligned. We also present a multi-modal alignment algorithm to train FaceCLIP, using a loss that aligns its joint representation with face, text, and image embedding spaces. We then build FaceCLIP-SDXL, an ID-preserving image synthesis pipeline by integrating FaceCLIP with Stable Diffusion XL (SDXL). Compared to prior methods, FaceCLIP-SDXL enables photorealistic portrait generation with better identity preservation and textual relevance. Extensive experiments demonstrate its quantitative and qualitative superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。