arXiv:2410.14247cs.CV2024-10被引 4

提出双链反演方法,实现高保真图像编辑与重建。

ERDDCI: Exact Reversible Diffusion via Dual-Chain Inversion for High-Quality Image Editing

  • 采用双链反演机制,实现可逆扩散过程。
  • 50步内重建SSIM达0.999,LPIPS仅0.001。
  • 适合高引导系数下的精细图像编辑任务。

扩散模型在真实图像编辑中已取得成功应用。现有方法通常将图像反演为用于重构原图的潜在噪声向量(即反演),并在推理过程中进行编辑。然而,当前主流扩散模型依赖局部线性化假设,即反演过程中注入的噪声需近似推理过程中移除的噪声。尽管该假设使生成高效,但会因累积误差影响图像重建与编辑质量。为此,本文提出新方法ERDDCI(基于双链反演的精确可逆扩散),通过双链反演(DCI)联合推断,实现精确可逆的扩散过程。该方法避免了传统反演中的复杂优化,显著提升图像编辑质量。此外,针对高引导系数下的图像操作,引入动态控制策略,实现更精细的重构与编辑。实验表明,ERDDCI在50步扩散过程中显著优于现有最优方法:重建SSIM达0.999,LPIPS为0.001,并在图像编辑任务中表现优异。

原文摘要 · Abstract (English)

Diffusion models (DMs) have been successfully applied to real image editing. These models typically invert images into latent noise vectors used to reconstruct the original images (known as inversion), and then edit them during the inference process. However, recent popular DMs often rely on the assumption of local linearization, where the noise injected during the inversion process is expected to approximate the noise removed during the inference process. While DM efficiently generates images under this assumption, it can also accumulate errors during the diffusion process due to the assumption, ultimately negatively impacting the quality of real image reconstruction and editing. To address this issue, we propose a novel method, referred to as ERDDCI (Exact Reversible Diffusion via Dual-Chain Inversion). ERDDCI uses the new Dual-Chain Inversion (DCI) for joint inference to derive an exact reversible diffusion process. By using DCI, our method effectively avoids the cumbersome optimization process in existing inversion approaches and achieves high-quality image editing. Additionally, to accommodate image operations under high guidance scales, we introduce a dynamic control strategy that enables more refined image reconstruction and editing. Our experiments demonstrate that ERDDCI significantly outperforms state-of-the-art methods in a 50-step diffusion process. It achieves rapid and precise image reconstruction with an SSIM of 0.999 and an LPIPS of 0.001, and also delivers competitive results in image editing.

图像编辑扩散模型反演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。