提出新方法逆推扩散模型生成初始噪声,精度显著提升。
Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel Method

- 先用梯度下降反推第一步,再用固定点法迭代后续步骤。
- 在三个数据集上重建精度和初始噪声预测均优于现有方法。
- 引入自插值测试,更真实评估生成图像质量,适合图像编辑研究者。
本文研究去噪扩散隐式模型(DDIM)生成过程的逆问题,旨在从生成图像中恢复潜在变量,特别是初始噪声图。现有方法在该任务中常表现不佳。我们提出一种混合新方法:首先通过梯度下降直接反推初始步骤,随后采用固定点法进行后续迭代。在三个数据集上的实证评估表明,该方法显著提升了初始潜在变量的预测精度,并实现了更优的重构性能。此外,我们引入新的自插值测试,评估在真实与预测潜变量之间插值点生成图像的质量,提供对模型性能的深入洞察。结果表明,尽管现有方法在重构上表现尚可,但始终无法准确预测初始潜变量,导致自插值测试表现差。而本方法在所有指标上均优于现有方法,为理解扩散模型提供了新视角,有助于提升图像生成与编辑应用。
原文摘要 · Abstract (English)
This paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a generated image. Existing methods often struggle with accuracy in this task. We propose a novel hybrid approach that combines direct inversion via gradient descent for the first step, followed by a fixed-point method for subsequent steps. Empirical evaluations across three datasets demonstrate that our method significantly improves the prediction of initial latent variables while achieving superior reconstruction accuracy. Additionally, we introduce a new evaluation, called the self-interpolation test, which assesses the quality of images generated from interpolated points between the true and predicted latent maps, offering deeper insights into performance. Our results reveal that while existing methods perform reasonably well in reconstruction, they consistently fail to accurately predict the initial latent variables, resulting in poor performance on the self-interpolation test. In contrast, our method outperforms all others across all metrics, providing valuable insights into diffusion models and enhancing their applications in image generation and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。