arXiv:2503.09277cs.CVcs.AI2025-03ICCV被引 31

统一处理多种条件的图像生成框架,支持文本、空间图、主体图任意组合。

UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer

  • 提出新型条件注意力机制与可训练LoRA模块,支持自由组合多条件输入。
  • 在SubjectSpatial200K数据集上实现最先进性能,多条件生成一致性显著提升。
  • 适合需要灵活可控生成的视觉生成研究者,尤其关注多条件融合场景。

随着扩散模型在图像生成中的快速发展,对更强大、灵活的可控生成框架的需求日益增长。尽管现有方法能超越文本提示进行引导生成,但如何有效组合多个条件输入并保持与所有条件的一致性仍是未解难题。为此,我们提出UniCombine,一种基于DiT的多条件可控生成框架,可处理任意组合的条件输入,包括但不限于文本提示、空间图和主体图像。具体地,我们引入一种新的条件MMDiT注意力机制,并结合可训练的LoRA模块,构建了无需训练和需训练两种版本。此外,我们提出新流程构建了首个专为多条件生成任务设计的数据集SubjectSpatial200K,涵盖主体驱动与空间对齐条件。在多条件生成任务上的大量实验表明,该方法具有卓越的通用性和强大生成能力,达到当前最优表现。

原文摘要 · Abstract (English)

With the rapid development of diffusion models in image generation, the demand for more powerful and flexible controllable frameworks is increasing. Although existing methods can guide generation beyond text prompts, the challenge of effectively combining multiple conditional inputs while maintaining consistency with all of them remains unsolved. To address this, we introduce UniCombine, a DiT-based multi-conditional controllable generative framework capable of handling any combination of conditions, including but not limited to text prompts, spatial maps, and subject images. Specifically, we introduce a novel Conditional MMDiT Attention mechanism and incorporate a trainable LoRA module to build both the training-free and training-based versions. Additionally, we propose a new pipeline to construct SubjectSpatial200K, the first dataset designed for multi-conditional generative tasks covering both the subject-driven and spatially-aligned conditions. Extensive experimental results on multi-conditional generation demonstrate the outstanding universality and powerful capability of our approach with state-of-the-art performance.

扩散模型可控生成多条件融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。