arXiv:2502.20376cs.GRcs.CV2025-02ICCV被引 9

用输入图像本身作为条件,提升真实图像编辑的重建与可编辑性。

Tight Inversion: Image-Conditioned Inversion for Real Image Editing

  • 以输入图像自身为精确条件进行反演,缩小模型输出分布。
  • 在高细节图像上重建精度提升,编辑效果更稳定。
  • 适配现有反演方法,适合追求高保真图像编辑的研究者。

文本到图像扩散模型具备强大的图像编辑能力。为编辑真实图像,许多方法依赖于将图像反演为高斯噪声。常见做法是逐步向图像添加噪声,其中噪声由逆向采样方程决定。该过程存在重建与可编辑性之间的固有权衡,限制了对高细节图像的编辑效果。鉴于文本到图像模型反演对文本条件的依赖,本文探究条件选择的重要性。结果表明,与输入图像高度对齐的条件能显著提升反演质量。基于此,我们提出紧致反演(Tight Inversion),采用最精确的条件——输入图像本身。该紧致条件缩小了模型输出分布,同时增强重建精度与可编辑性。通过大量实验验证,本方法在结合现有反演技术时表现优异,评估涵盖重建准确率及与多种编辑方法的兼容性。

原文摘要 · Abstract (English)

Text-to-image diffusion models offer powerful image editing capabilities. To edit real images, many methods rely on the inversion of the image into Gaussian noise. A common approach to invert an image is to gradually add noise to the image, where the noise is determined by reversing the sampling equation. This process has an inherent tradeoff between reconstruction and editability, limiting the editing of challenging images such as highly-detailed ones. Recognizing the reliance of text-to-image models inversion on a text condition, this work explores the importance of the condition choice. We show that a condition that precisely aligns with the input image significantly improves the inversion quality. Based on our findings, we introduce Tight Inversion, an inversion method that utilizes the most possible precise condition -- the input image itself. This tight condition narrows the distribution of the model's output and enhances both reconstruction and editability. We demonstrate the effectiveness of our approach when combined with existing inversion methods through extensive experiments, evaluating the reconstruction accuracy as well as the integration with various editing methods.

图像编辑扩散模型反演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。