无需训练调优即可精准还原真实图像并实现语义一致的编辑
Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
- 提出双调度反演方法,数学上保证图像可逆性
- 无需微调即可完美重建真实图像,编辑部分符合文本语义
- 适合需要快速、稳定图像编辑的开发者和设计师
文本条件图像编辑是近年来具有重要商业与学术价值的AIGC任务。针对真实图像编辑,多数基于扩散模型的方法在编辑前使用DDIM反演作为第一阶段,但该方法常导致重建失败,影响下游编辑效果。本文首先分析了DDIM反演重建失败的原因,并提出一种新的反演与采样方法——双调度反演(Dual-Schedule Inversion)。同时设计了一个分类器,可自适应地将双调度反演与不同编辑方法结合,提升用户体验。本工作在无需训练或微调的情况下,实现了卓越的重建与编辑性能:1)可完美重建真实图像,且其可逆性在数学上得到保证;2)编辑内容严格遵循文本提示的语义;3)未编辑区域保留原始身份特征。
原文摘要 · Abstract (English)
Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。