无需训练即可融合多种条件生成图像,突破单条件限制。
MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis

- 通过理论推导的整合公式融合多个预训练模型的噪声预测
- 支持任意数量的控制条件,包括文本、关键点、草图等
- 无需重新训练,可快速扩展新条件模态,适合快速原型开发
近期文本到图像(T2I)生成进展实现了通过超越文本的条件进行可控图像合成。然而,现有基于扩散的方法大多仅支持单一类型的控制条件(如边界框或关键点),限制了灵活性。为此,我们提出 MixDiffusion,一种无需训练的多条件 T2I 生成扩散框架。该方法理论上支持任意数量的控制条件,包括边界框、关键点、草图、深度图、参考图像和文本,通过协同集成多个预训练的单条件扩散模型实现。核心思路是:在每一步去噪过程中,从多个单条件扩散模型的噪声预测分布中,通过推导出的整合公式得到最终的噪声分布,且该方法具备严格的理论证明。由于无需训练,MixDiffusion 易于部署,并可轻松扩展至新的控制模态。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have enabled controllable image synthesis by incorporating conditions beyond text. However, most existing diffusion-based methods are limited to a single type of control condition (e.g., bounding boxes or keypoints), which restricts their flexibility. To address this limitation, we propose MixDiffusion, a training-free diffusion framework for multi-condition T2I generation. MixDiffusion theoretically supports an arbitrary number of control conditions, including bounding boxes, keypoints, sketches, depth maps, reference images, and text, by collaboratively integrating multiple pre-trained uni-condition diffusion models. The key insight of the proposed approach is to derive the predicted noise distribution in each denoising step of the diffusion-based multi-condition image generation model from the predicted noise distributions of multiple diffusion-based uni-condition models with a derived integration formula, which is supported by rigorous theory proof. Owing to its training-free nature, MixDiffusion is easy to deploy and readily extensible to new control modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。