arXiv:2410.10792cs.LGcs.CV2024-10ICLR被引 177

用随机微分方程实现图像逆向生成与编辑,无需额外训练。

Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations

  • 基于线性二次调节器设计动态最优控制,实现矩形流模型的逆向生成。
  • 零样本逆向重建与语义编辑性能达到当前最佳,人眼评估更受青睐。
  • 方法无需训练参数或测试时优化,适用于快速图像编辑任务。

生成模型将随机噪声转换为图像;其逆向过程旨在将图像还原为结构化噪声以实现恢复与编辑。本文针对两大核心任务:(i) 使用矩形流模型(如 Flux)的随机等价形式进行图像逆向,(ii) 实现真实图像的编辑。尽管扩散模型(DMs)近年来主导图像生成领域,但其逆向因漂移和扩散项的非线性导致忠实度与可编辑性挑战。现有先进方法依赖额外参数训练或测试时隐变量优化,实际应用成本高。矩形流(RFs)是扩散模型的有前景替代方案,但其逆向研究仍不充分。本文提出基于线性二次调节器推导的动态最优控制进行矩形流逆向,证明所得向量场等价于矩形随机微分方程。此外,我们扩展框架设计了 Flux 的随机采样器。该逆向方法在零样本逆向与编辑上达到顶尖性能,在笔触到图像合成与语义图像编辑方面超越先前工作,大规模人工评估确认用户偏好。

原文摘要 · Abstract (English)

Generative models transform random noise into images; their inversion aims to transform images back to structured noise for recovery and editing. This paper addresses two key tasks: (i) inversion and (ii) editing of a real image using stochastic equivalents of rectified flow models (such as Flux). Although Diffusion Models (DMs) have recently dominated the field of generative modeling for images, their inversion presents faithfulness and editability challenges due to nonlinearities in drift and diffusion. Existing state-of-the-art DM inversion approaches rely on training of additional parameters or test-time optimization of latent variables; both are expensive in practice. Rectified Flows (RFs) offer a promising alternative to diffusion models, yet their inversion has been underexplored. We propose RF inversion using dynamic optimal control derived via a linear quadratic regulator. We prove that the resulting vector field is equivalent to a rectified stochastic differential equation. Additionally, we extend our framework to design a stochastic sampler for Flux. Our inversion method allows for state-of-the-art performance in zero-shot inversion and editing, outperforming prior works in stroke-to-image synthesis and semantic image editing, with large-scale human evaluations confirming user preference.

图像逆向生成模型扩散模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。