arXiv:2409.10476cs.CV2024-09被引 2

改进扩散模型图像编辑中的反演误差,提升精度且不降效率。

SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing

  • 分离源图与目标图的引导尺度,降低反演误差
  • 理论推导出更优引导尺度为0.5,性能显著提升
  • 保持原框架,适合需高精度文本图像编辑的研究者

扩散模型在文本引导下展现出强大的图像生成能力。受扩散学习过程启发,现有方法可通过DDIM反演实现基于文本的图像编辑。然而,原始的DDIM反演未针对无分类器引导优化,累积误差导致性能下降。尽管已有诸多算法改进该框架,本文研究了DDIM反演中的近似误差,提出将源分支与目标分支的引导尺度解耦,以在保持原有框架的前提下减少误差。此外,理论推导表明更优的引导尺度为0.5。在PIE-Bench上的实验显示,本方法可显著提升DDIM反演性能,且不牺牲效率。

原文摘要 · Abstract (English)

Diffusion models demonstrate impressive image generation performance with text guidance. Inspired by the learning process of diffusion, existing images can be edited according to text by DDIM inversion. However, the vanilla DDIM inversion is not optimized for classifier-free guidance and the accumulated error will result in the undesired performance. While many algorithms are developed to improve the framework of DDIM inversion for editing, in this work, we investigate the approximation error in DDIM inversion and propose to disentangle the guidance scale for the source and target branches to reduce the error while keeping the original framework. Moreover, a better guidance scale (i.e., 0.5) than default settings can be derived theoretically. Experiments on PIE-Bench show that our proposal can improve the performance of DDIM inversion dramatically without sacrificing efficiency.

图像编辑扩散模型反演文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。