动态适配器提升零样本个性化图像生成的细节与多主体扩展性
DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- 通过动态解耦策略分离概念无关信息,平衡概念保留与提示遵循
- 引入分层专家融合模块,显著增强参考图细粒度特征保留能力
- 支持多主体个性化生成,适用于需要高保真定制图像的场景
个性化文本到图像(PT2I)生成旨在基于参考图像生成定制化图像。当前方法在零样本场景下面临三大挑战:概念保留(CP)与提示遵循(PF)难以平衡、参考图像中的细粒度概念难以保留、难以扩展至多主体个性化。为此,本文提出动态图像提示适配器(DynaIP),作为SOTA文本到图像多模态扩散模型(MM-DiT)的插件,提升细粒度概念保真度、CP-PF平衡性和主体可扩展性。关键发现是:当通过交叉注意力注入参考图像特征时,MM-DiT呈现内在解耦学习行为。基于此,设计动态解耦策略,在推理阶段去除概念无关信息干扰,显著改善平衡性并提升多主体组合可扩展性。此外,识别出视觉编码器是影响细粒度概念保留的关键因素,发现常用CLIP的分层特征可捕捉多粒度视觉信息。因此提出分层专家融合模块,充分利用CLIP的分层特征,显著提升细粒度保真度,并实现视觉粒度灵活控制。单主体与多主体实验验证表明,DynaIP优于现有方法,推动了该领域的进展。
原文摘要 · Abstract (English)
Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images. A prominent interest pertains to the integration of an image prompt adapter to facilitate zero-shot PT2I without test-time fine-tuning. However, current methods grapple with three fundamental challenges: 1. the elusive equilibrium between Concept Preservation (CP) and Prompt Following (PF), 2. the difficulty in retaining fine-grained concept details in reference images, and 3. the restricted scalability to extend to multi-subject personalization. To tackle these challenges, we present Dynamic Image Prompt Adapter (DynaIP), a cutting-edge plugin to enhance the fine-grained concept fidelity, CP-PF balance, and subject scalability of SOTA T2I multimodal diffusion transformers (MM-DiT) for PT2I generation. Our key finding is that MM-DiT inherently exhibit decoupling learning behavior when injecting reference image features into its dual branches via cross attentions. Based on this, we design an innovative Dynamic Decoupling Strategy that removes the interference of concept-agnostic information during inference, significantly enhancing the CP-PF balance and further bolstering the scalability of multi-subject compositions. Moreover, we identify the visual encoder as a key factor affecting fine-grained CP and reveal that the hierarchical features of commonly used CLIP can capture visual information at diverse granularity levels. Therefore, we introduce a novel Hierarchical Mixture-of-Experts Feature Fusion Module to fully leverage the hierarchical features of CLIP, remarkably elevating the fine-grained concept fidelity while also providing flexible control of visual granularity. Extensive experiments across single- and multi-subject PT2I tasks verify that our DynaIP outperforms existing approaches, marking a notable advancement in the field of PT2l generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。