arXiv:2505.22360cs.CV2025-05被引 2

解决文本生成图像时人物身份保持难题,提升定制化图像质量。

Identity-Preserving Text-to-Image Generation via Dual-Level Feature Decoupling and Expert-Guided Fusion

  • 通过双层特征解耦分离人物身份与背景信息
  • 采用专家混合机制动态融合特征,提升生成多样性
  • 适合需要精准身份保留的个性化图像生成场景

大规模文本到图像生成模型推动了以主体为中心的图像生成发展,旨在根据文本描述生成符合要求的定制图像并保持特定主体的身份。尽管取得显著进展,现有方法难以有效分离输入图像中与身份相关和无关的信息,导致过拟合或身份丢失。本文提出一种新框架,通过改进身份相关与无关特征的分离,并引入创新的特征融合机制,提升生成图像的质量与文本对齐度。框架包含两个关键组件:隐式-显式前景-背景解耦模块(IEDM)和基于专家混合(MoE)的特征融合模块(FFM)。IEDM结合可学习适配器实现特征级隐式解耦,以及基于修复技术的图像级显式分离。FFM动态整合身份无关特征与身份相关特征,即使在解耦不完全的情况下也能生成更优特征表示。此外,引入三种互补损失函数引导解耦过程。大量实验表明,该方法能显著提升生成质量、增强场景适应灵活性,并在不同文本描述下增加输出多样性。

原文摘要 · Abstract (English)

Recent advances in large-scale text-to-image generation models have led to a surge in subject-driven text-to-image generation, which aims to produce customized images that align with textual descriptions while preserving the identity of specific subjects. Despite significant progress, current methods struggle to disentangle identity-relevant information from identity-irrelevant details in the input images, resulting in overfitting or failure to maintain subject identity. In this work, we propose a novel framework that improves the separation of identity-related and identity-unrelated features and introduces an innovative feature fusion mechanism to improve the quality and text alignment of generated images. Our framework consists of two key components: an Implicit-Explicit foreground-background Decoupling Module (IEDM) and a Feature Fusion Module (FFM) based on a Mixture of Experts (MoE). IEDM combines learnable adapters for implicit decoupling at the feature level with inpainting techniques for explicit foreground-background separation at the image level. FFM dynamically integrates identity-irrelevant features with identity-related features, enabling refined feature representations even in cases of incomplete decoupling. In addition, we introduce three complementary loss functions to guide the decoupling process. Extensive experiments demonstrate the effectiveness of our proposed method in enhancing image generation quality, improving flexibility in scene adaptation, and increasing the diversity of generated outputs across various textual descriptions.

文本生成图像身份保持特征解耦MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。