提出显式条件机制,让图像编辑更快更高效。
Guidance Free Image Editing via Explicit Conditioning
- 直接建模噪声分布以实现条件控制
- 仅需一次去噪即可生成高质量图像
- 适合追求高效推理的图像生成应用
当前基于条件扩散模型的采样方法主要依赖无分类器引导(CFG)生成高质量图像,但其在每一步需多次去噪,例如图像编辑任务中最多需三次,导致计算成本过高。本文提出一种新型条件机制——显式条件(EC),通过将噪声分布显式地与输入模态关联,减轻了传统引导技术的计算负担,显著提升扩散模型的推理速度。实验表明,在图像编辑任务中,该方法在显著降低计算量的同时,仍能生成多样且高质量的图像,性能优于现有CFG方案。
原文摘要 · Abstract (English)
Current sampling mechanisms for conditional diffusion models rely mainly on Classifier Free Guidance (CFG) to generate high-quality images. However, CFG requires several denoising passes in each time step, e.g., up to three passes in image editing tasks, resulting in excessive computational costs. This paper introduces a novel conditioning technique to ease the computational burden of the well-established guidance techniques, thereby significantly improving the inference time of diffusion models. We present Explicit Conditioning (EC) of the noise distribution on the input modalities to achieve this. Intuitively, we model the noise to guide the conditional diffusion model during the diffusion process. We present evaluations on image editing tasks and demonstrate that EC outperforms CFG in generating diverse high-quality images with significantly reduced computations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。