让多个扩散模型像团队协作一样生成复杂图像,通过最优控制协调轨迹。
CMAD: Cooperative Multi-Agent Diffusion via Stochastic Optimal Control
- 将多个预训练扩散模型视为协作智能体,用最优控制联合引导生成轨迹。
- 在条件MNIST生成任务上优于基线方法,实现更优的图像组合质量。
- 适合需要多模型协同生成的场景,如复杂图像合成与可控设计。
连续时间生成模型在图像修复与合成方面取得了显著进展。然而,如何控制多个预训练模型的组合仍是一个开放问题。现有方法大多将组合视为概率密度的代数组合,如乘积或专家混合,这假设目标分布显式已知,但实际几乎从未满足。本文提出一种新范式:将组合生成建模为合作随机最优控制问题。不直接组合概率密度,而是将预训练扩散模型视为相互作用的智能体,通过最优控制联合引导其扩散轨迹,使其在聚合输出上趋近共享目标。我们在条件MNIST生成任务上验证了该框架,并与一种朴素的推理时基线方法对比,该基线以每步梯度引导替代学习到的合作控制。结果表明,本方法在生成质量上表现更优。
原文摘要 · Abstract (English)
Continuous-time generative models have achieved remarkable success in image restoration and synthesis. However, controlling the composition of multiple pre-trained models remains an open challenge. Current approaches largely treat composition as an algebraic composition of probability densities, such as via products or mixtures of experts. This perspective assumes the target distribution is known explicitly, which is almost never the case. In this work, we propose a different paradigm that formulates compositional generation as a cooperative Stochastic Optimal Control problem. Rather than combining probability densities, we treat pre-trained diffusion models as interacting agents whose diffusion trajectories are jointly steered, via optimal control, toward a shared objective defined on their aggregated output. We validate our framework on conditional MNIST generation and compare it against a naïve inference-time DPS-style baseline replacing learned cooperative control with per-step gradient guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。