发现大模型反演时的振荡现象,揭示其结构意义并用于图像增强与编辑
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
- 用迭代反演法观察到图像反演不收敛,而是振荡于不同语义簇
- 振荡簇具语义一致性,且可实现高效图像增强与笔触着色等操作
- 方法简单快速,适合需要可控图像编辑的开发者和研究者
我们研究了在大规模文本到图像扩散模型(以'Flux'模型为例)中应用反演方法时观察到的振荡行为。通过采用受固定点启发的迭代反演方法处理真实图像,发现解无法收敛,而是持续在不同语义簇间振荡。通过简化实验和真实扩散模型验证,这些振荡簇展现出显著的语义一致性。理论分析表明,该现象源于修正流模型中的振荡动力学。基于此理解,我们提出一种简单高效的分布迁移技术,可实现图像增强、基于笔触的重新着色及视觉提示引导的图像编辑。定量结果表明,该方法在图像增强、妆容迁移、重建质量与引导采样质量等方面均表现优异。更多高质量图像与视频示例见:https://yanyanzheng96.github.io/oscillation_inversion/
原文摘要 · Abstract (English)
We explore the oscillatory behavior observed in inversion methods applied to large-scale text-to-image diffusion models, with a focus on the "Flux" model. By employing a fixed-point-inspired iterative approach to invert real-world images, we observe that the solution does not achieve convergence, instead oscillating between distinct clusters. Through both toy experiments and real-world diffusion models, we demonstrate that these oscillating clusters exhibit notable semantic coherence. We offer theoretical insights, showing that this behavior arises from oscillatory dynamics in rectified flow models. Building on this understanding, we introduce a simple and fast distribution transfer technique that facilitates image enhancement, stroke-based recoloring, as well as visual prompt-guided image editing. Furthermore, we provide quantitative results demonstrating the effectiveness of our method for tasks such as image enhancement, makeup transfer, reconstruction quality, and guided sampling quality. Higher-quality examples of videos and images are available at \href{https://yanyanzheng96.github.io/oscillation_inversion/}{this link}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。