改进生成模型反演稳定性,无需训练即可提升图像重建与编辑质量。
Free Lunch for Stabilizing Rectified Flow Inversion
- 通过运行平均速度引导反演路径,稳定梯度更新过程。
- 在PIE-Bench上实现更优重建与编辑效果,减少函数求值次数。
- 适合需要高效、高保真图像重构与编辑的研究者使用。
基于修正流(Rectified-Flow, RF)的生成模型已成为传统扩散模型的有力替代,已在多项任务中达到领先性能。其通过学习连续速度场将简单噪声转化为复杂数据,不仅支持高质量生成,还具备无需训练的反演能力,适用于重建与编辑等下游任务。然而,现有反演方法(如原始RF反演)存在随时间步累积的近似误差,导致速度场不稳定,降低重建与编辑质量。为此,我们提出近端均值反演(Proximal-Mean Inversion, PMI),一种无需训练的梯度修正方法,通过将速度场约束在理论推导出的球形高斯内,向历史速度的运行平均值引导,以稳定反演过程。此外,我们引入mimic-CFG,一种轻量级速度修正方案,用于编辑任务中在当前速度与历史平均投影间插值,平衡编辑效果与结构一致性。在PIE-Bench上的大量实验表明,我们的方法显著提升了反演稳定性、图像重建质量与编辑保真度,同时减少神经函数评估次数。该方法在PIE-Bench上实现了最先进的性能,兼具效率与理论严谨性。
原文摘要 · Abstract (English)
Rectified-Flow (RF)-based generative models have recently emerged as strong alternatives to traditional diffusion models, demonstrating state-of-the-art performance across various tasks. By learning a continuous velocity field that transforms simple noise into complex data, RF-based models not only enable high-quality generation, but also support training-free inversion, which facilitates downstream tasks such as reconstruction and editing. However, existing inversion methods, such as vanilla RF-based inversion, suffer from approximation errors that accumulate across timesteps, leading to unstable velocity fields and degraded reconstruction and editing quality. To address this challenge, we propose Proximal-Mean Inversion (PMI), a training-free gradient correction method that stabilizes the velocity field by guiding it toward a running average of past velocities, constrained within a theoretically derived spherical Gaussian. Furthermore, we introduce mimic-CFG, a lightweight velocity correction scheme for editing tasks, which interpolates between the current velocity and its projection onto the historical average, balancing editing effectiveness and structural consistency. Extensive experiments on PIE-Bench demonstrate that our methods significantly improve inversion stability, image reconstruction quality, and editing fidelity, while reducing the required number of neural function evaluations. Our approach achieves state-of-the-art performance on the PIE-Bench with enhanced efficiency and theoretical soundness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。