arXiv:2410.04844cs.CVcs.AI2024-10ICLR被引 10

无需优化和训练,1.5秒实现高效且背景一致的零样本图像编辑。

PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing

  • 通过后验采样控制扩散过程,结合初始特征与Langevin动力学
  • 生成结果在未编辑区域保持高度一致性,性能达当前最优
  • 无需反演或训练,仅需1.5秒和18GB显存,适合快速部署

图像编辑领域仍面临可控性、背景保持和效率三大挑战。基于反演的方法依赖耗时优化以保留初始图像特征,效率低下;而无反演方法缺乏理论支撑,难以保证背景相似性。为此,我们提出PostEdit,引入后验机制调控扩散采样过程。具体地,设计一个关联初始特征与Langevin动力学的测量项,优化由目标提示生成的估计图像。大量实验表明,PostEdit在保持未编辑区域一致性的同时达到最先进的编辑性能。该方法无需反演也无需训练,生成高质量结果仅需约1.5秒及18 GB GPU内存。

原文摘要 · Abstract (English)

In the field of image editing, three core challenges persist: controllability, background preservation, and efficiency. Inversion-based methods rely on time-consuming optimization to preserve the features of the initial images, which results in low efficiency due to the requirement for extensive network inference. Conversely, inversion-free methods lack theoretical support for background similarity, as they circumvent the issue of maintaining initial features to achieve efficiency. As a consequence, none of these methods can achieve both high efficiency and background consistency. To tackle the challenges and the aforementioned disadvantages, we introduce PostEdit, a method that incorporates a posterior scheme to govern the diffusion sampling process. Specifically, a corresponding measurement term related to both the initial features and Langevin dynamics is introduced to optimize the estimated image generated by the given target prompt. Extensive experimental results indicate that the proposed PostEdit achieves state-of-the-art editing performance while accurately preserving unedited regions. Furthermore, the method is both inversion- and training-free, necessitating approximately 1.5 seconds and 18 GB of GPU memory to generate high-quality results.

图像编辑扩散模型零样本高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。