改进扩散模型噪声调度,让图像编辑更保真
Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
- 提出新型逻辑噪声调度,解决传统调度的奇异点问题
- 在8项编辑任务中显著提升内容保留率和编辑保真度
- 无需重训练,可通用适配现有编辑方法,适合图像生成研究者
文本引导的扩散模型已大幅推动图像编辑技术,实现高质量、多样化的文本驱动修改。然而,有效编辑需将源图像反演至隐空间,此过程常受DDIM反演中的预测误差影响。这些误差在扩散过程中累积,导致内容保留与编辑保真度下降,尤其在条件输入下更为明显。本文通过分析DDIM反演中误差累积的主要来源,识别出传统噪声调度的奇异点问题是关键症结。为此,提出新型逻辑噪声调度(Logistic Schedule),旨在消除奇异点、提升反演稳定性,并提供更优的噪声空间用于图像编辑。该调度降低噪声预测误差,实现更忠实的编辑效果,同时保持源图像原有内容。本方法无需额外训练,兼容多种现有编辑方法。在8项编辑任务上的实验表明,其在内容保留与编辑保真度方面优于传统噪声调度,凸显其适应性与有效性。
原文摘要 · Abstract (English)
Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion. These errors accumulate during the diffusion process, resulting in inferior content preservation and edit fidelity, especially with conditional inputs. We address these challenges by investigating the primary contributors to error accumulation in DDIM inversion and identify the singularity problem in traditional noise schedules as a key issue. To resolve this, we introduce the Logistic Schedule, a novel noise schedule designed to eliminate singularities, improve inversion stability, and provide a better noise space for image editing. This schedule reduces noise prediction errors, enabling more faithful editing that preserves the original content of the source image. Our approach requires no additional retraining and is compatible with various existing editing methods. Experiments across eight editing tasks demonstrate the Logistic Schedule's superior performance in content preservation and edit fidelity compared to traditional noise schedules, highlighting its adaptability and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。