arXiv:2412.15798cs.CV2024-12被引 6

无需训练,用提示词精准修改图像结构与内容。

Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

  • 通过双目标引导优化反向过程潜变量,实现结构保留与提示对齐。
  • 在多个任务上表现优异,生成图像与目标提示匹配度高。
  • 适合需要快速图像编辑且不希望重新训练的用户使用。

我们提出一种基于预训练文本到图像扩散模型的简单而有效的无训练文本驱动图像到图像转换方法。目标是在保持源图像结构和背景的前提下,生成符合目标任务的图像。为此,我们通过结合两个目标推导出表示引导:最大化基于CLIP分数的目标提示相似性,同时最小化源潜在变量的结构距离。该引导提升了生成目标图像对目标提示的保真度,同时保持了源图像的结构完整性。为引入表示引导,我们对扩散模型反向过程的目标潜在变量进行优化。实验结果表明,该方法在结合预训练Stable Diffusion模型时,在多种任务上均实现了出色的图像到图像转换性能。

原文摘要 · Abstract (English)

We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model.

图像编辑扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。