系统梳理扩散模型的条件图像生成方法,助你快速掌握核心思路与应用。
Conditional Image Synthesis with Diffusion Models: A Survey
- 按条件如何融入去噪网络和采样过程分类,理清技术脉络
- 总结六种主流采样阶段条件机制,覆盖常见应用场景
- 适合想入门或跟进扩散模型研究的研究者参考
基于用户指定要求的条件图像合成是构建复杂视觉内容的关键。近年来,基于扩散的生成建模已成为条件图像合成的有效手段,相关文献呈指数级增长。然而,扩散建模的复杂性、图像合成任务的多样性以及条件机制的差异性,给研究人员跟踪进展和理解核心概念带来挑战。本文根据条件如何融入扩散建模的两个基础组件——去噪网络和采样过程——对现有工作进行分类。重点分析不同条件方法在训练、再利用和专业化阶段的底层原理、优势与潜在挑战,以构建理想的去噪网络。同时总结了六种主流的采样过程条件机制。所有讨论均围绕典型应用展开。最后,指出了若干关键但尚未解决的问题,并提出未来研究可能的解决方案。相关工作列表见:https://github.com/zju-pi/Awesome-Conditional-Diffusion-Models。
原文摘要 · Abstract (English)
Conditional image synthesis based on user-specified requirements is a key component in creating complex visual content. In recent years, diffusion-based generative modeling has become a highly effective way for conditional image synthesis, leading to exponential growth in the literature. However, the complexity of diffusion-based modeling, the wide range of image synthesis tasks, and the diversity of conditioning mechanisms present significant challenges for researchers to keep up with rapid developments and to understand the core concepts on this topic. In this survey, we categorize existing works based on how conditions are integrated into the two fundamental components of diffusion-based modeling, $\textit{i.e.}$, the denoising network and the sampling process. We specifically highlight the underlying principles, advantages, and potential challenges of various conditioning approaches during the training, re-purposing, and specialization stages to construct a desired denoising network. We also summarize six mainstream conditioning mechanisms in the sampling process. All discussions are centered around popular applications. Finally, we pinpoint several critical yet still unsolved problems and suggest some possible solutions for future research. Our reviewed works are itemized at https://github.com/zju-pi/Awesome-Conditional-Diffusion-Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。