通过上下文生成实现多主体图像可控生成,解决数据扩展难题
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

- 利用扩散变换器的上下文生成能力合成高一致性多主体数据
- 在单/多主体场景下均实现高一致性和强可控性
- 适合需要多主体图像生成与控制的研究者
尽管基于主体的图像生成在实际应用中已广泛研究,但在数据可扩展性和主体泛化性方面仍面临挑战。首先,从单主体数据集扩展到多主体数据集并进行规模化构建极为困难;其次,现有方法多集中于单主体生成,难以应对多主体场景。为此,本文提出一种高度一致的数据合成流程,利用扩散变换器的内在上下文生成能力,生成高一致性多主体配对数据。此外,我们引入UNO模型,包含渐进式跨模态对齐和通用旋转位置编码,是一个从文本到图像模型迭代训练得到的多图像条件主体到图像模型。大量实验表明,该方法在单主体和多主体驱动生成中均能实现高一致性与良好可控性。
原文摘要 · Abstract (English)
Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating single-subject datasets to multiple-subject ones and scaling them is particularly difficult. For the second, most recent methods center on single-subject generation, making it hard to apply when dealing with multi-subject scenarios. In this study, we propose a highly-consistent data synthesis pipeline to tackle this challenge. This pipeline harnesses the intrinsic in-context generation capabilities of diffusion transformers and generates high-consistency multi-subject paired data. Additionally, we introduce UNO, which consists of progressive cross-modal alignment and universal rotary position embedding. It is a multi-image conditioned subject-to-image model iteratively trained from a text-to-image model. Extensive experiments show that our method can achieve high consistency while ensuring controllability in both single-subject and multi-subject driven generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。