arXiv:2507.11990cs.CV2025-07

让AI生成人脸更像本人,解决身份模糊问题。

ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation

  • 用身份嵌入引导文本特征,对齐图文语义
  • 生成速度比现有方法快15倍,身份保留更好
  • 适合个性化人像生成、数字肖像创作场景

近期,基于文本到图像扩散模型的个性化人像生成取得显著进展,文本反转(Textual Inversion)已成为生成高保真个性化图像的有前景方法。然而,当前文本反转方法因文本与视觉嵌入空间在身份语义上存在错位,难以保持一致的人脸身份。本文提出ID-EA框架,通过引导文本嵌入与视觉身份嵌入对齐,提升个性化生成中的身份保留能力。ID-EA包含两个核心组件:身份驱动增强器(ID-Enhancer)和身份条件适配器(ID-Adapter)。首先,ID-Enhancer将身份嵌入与文本身份锚点结合,利用代表性文本嵌入优化来自人脸识别模型的视觉身份嵌入;随后,ID-Adapter利用增强后的身份嵌入适配文本条件,通过调整预训练UNet模型中的交叉注意力模块,确保身份一致性,并促使文本特征在前景片段中寻找最相关视觉线索。大量定量与定性评估表明,ID-EA在身份保留指标上显著优于当前最优方法,同时实现卓越的计算效率,个性化人像生成速度约为现有方法的15倍。

原文摘要 · Abstract (English)

Recently, personalized portrait generation with a text-to-image diffusion model has significantly advanced with Textual Inversion, emerging as a promising approach for creating high-fidelity personalized images. Despite its potential, current Textual Inversion methods struggle to maintain consistent facial identity due to semantic misalignments between textual and visual embedding spaces regarding identity. We introduce ID-EA, a novel framework that guides text embeddings to align with visual identity embeddings, thereby improving identity preservation in a personalized generation. ID-EA comprises two key components: the ID-driven Enhancer (ID-Enhancer) and the ID-conditioned Adapter (ID-Adapter). First, the ID-Enhancer integrates identity embeddings with a textual ID anchor, refining visual identity embeddings derived from a face recognition model using representative text embeddings. Then, the ID-Adapter leverages the identity-enhanced embedding to adapt the text condition, ensuring identity preservation by adjusting the cross-attention module in the pre-trained UNet model. This process encourages the text features to find the most related visual clues across the foreground snippets. Extensive quantitative and qualitative evaluations demonstrate that ID-EA substantially outperforms state-of-the-art methods in identity preservation metrics while achieving remarkable computational efficiency, generating personalized portraits approximately 15 times faster than existing approaches.

个性化生成文本反转人脸生成身份保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。