让预训练扩散模型学会生成更大更复杂的图像和音频。
Diffusion Domain Expansion: Learning to Coordinate Pre-trained Diffusion Models

- 用轻量协调网络控制多个预训练扩散模型输出。
- 可生成比训练时更大范围的图像与音频,效果优于现有方法。
- 适合想扩展扩散模型能力的研究者和开发者。
本文提出扩散域扩展(DDE),一种高效扩展预训练扩散模型生成能力的方法,使其能够生成更大对象并处理更复杂的条件输入。该方法设计了一个轻量级可训练网络,用于协调多个预训练扩散模型的去噪输出。实验表明,该协调器虽结构简单,却能泛化到训练阶段未覆盖的更大领域。我们在长音频生成和条件图像生成任务上评估了DDE,验证了其跨领域的适用性。在定性和定量评价中,DDE均优于其他基于扩散模型的协同生成方法。
原文摘要 · Abstract (English)
In this paper, we propose Diffusion Domain Expansion (DDE), a method that efficiently extends pre-trained diffusion models to generate larger objects and handle more complex conditioning beyond their original capabilities. Our method employs a compact trainable network designed to coordinate the denoised outputs of pre-trained diffusion models. We demonstrate that the coordinator can be universally simple while being capable of generalizing to domains larger than those observed during its training time. We evaluate DDE on long audio track generation and conditional image generation, demonstrating its applicability across domains. DDE outperforms other approaches to coordinated generation with diffusion models in qualitative and quantitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。