无需微调即可个性化生成物体和抽象概念的图像。
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
- 通过调制适配器预测概念专属调制方向,实现无微调个性化。
- 在包含抽象概念的新基准上达到最优性能,生成效果更自然。
- 适合需要快速、稳定个性化生成的开发者与设计师使用。
个性化文本到图像生成旨在合成用户提供的概念在多样化场景中的图像。尽管多概念个性化取得进展,但多数方法仅限于物体概念,难以定制抽象概念(如姿态、光照)。部分方法开始支持抽象概念,但需为每个新概念进行测试时微调,耗时且易过拟合。本文提出一种全新的无微调多概念个性化方法,可有效定制物体与抽象概念。该方法基于预训练扩散变压器(DiT)模型中的调制机制,利用调制空间的局部化与语义意义。我们提出一种新模块——调制适配器(Mod-Adapter),用于为相关文本标记的调制过程预测概念特异性调制方向。该模块引入视觉-语言交叉注意力提取概念视觉特征,并采用混合专家(MoE)层将特征自适应映射至调制空间。为缓解概念图像空间与调制空间间巨大差距带来的训练困难,我们设计了基于视觉-语言模型的预训练策略,利用其强大的图像理解能力提供语义监督信号。为全面评估,我们在标准基准上扩展了抽象概念。实验表明,本方法在多概念个性化任务中达到当前最优性能,经定量、定性和人类评估验证。
原文摘要 · Abstract (English)
Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to customize abstract concepts (e.g., pose, lighting). Some methods have begun exploring multi-concept personalization supporting abstract concepts, but they require test-time fine-tuning for each new concept, which is time-consuming and prone to overfitting on limited training images. In this work, we propose a novel tuning-free method for multi-concept personalization that can effectively customize both object and abstract concepts without test-time fine-tuning. Our method builds upon the modulation mechanism in pre-trained Diffusion Transformers (DiTs) model, leveraging the localized and semantically meaningful properties of the modulation space. Specifically, we propose a novel module, Mod-Adapter, to predict concept-specific modulation direction for the modulation process of concept-related text tokens. It introduces vision-language cross-attention for extracting concept visual features, and Mixture-of-Experts (MoE) layers that adaptively map the concept features into the modulation space. Furthermore, to mitigate the training difficulty caused by the large gap between the concept image space and the modulation space, we introduce a VLM-guided pre-training strategy that leverages the strong image understanding capabilities of vision-language models to provide semantic supervision signals. For a comprehensive comparison, we extend a standard benchmark by incorporating abstract concepts. Our method achieves state-of-the-art performance in multi-concept personalization, supported by quantitative, qualitative, and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。