无需训练,用多步反演实现精准图像编辑。
Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness

- 采用多历史步预测-校正框架,减少迭代次数提升效率。
- 在30个细粒度任务上优于现有方法,保持未编辑区域准确。
- 自动生成语义角度掩码,支持大范围编辑且无需用户输入。
近年来,基于扩散随机微分方程(SDE)反演和无反演的方法成为免训练图像编辑主流,可在不微调模型的情况下实现高保真重建。然而,现有方法效率低、可塑性差,难以准确保留未编辑区域。为此,我们提出MIEdit,一种基于SDE反演的免训练编辑框架。MIEdit引入预测-校正多历史步机制,在更少步骤下实现更优编辑质量;通过缓解采样过程中多条件噪声残差与梯度项间的异质性与冲突,提升大范围编辑下的稳定性与可塑性。此外,框架包含反演时自动语义角度掩码(IASM),利用无分类器引导在反演阶段自动生成语义掩码,并贯穿整个采样过程施加区域约束,无需额外用户输入。我们还构建了EditEval++(含30个细粒度任务,1,000+图像-文本-掩码三元组)用于全面评估;实验表明,MIEdit显著优于当前最优技术。
原文摘要 · Abstract (English)
In recent years, diffusion stochastic differential equation (SDE) inversion and inversion-free methods have become prevalent for training-free image editing, as they can achieve faithful reconstruction without tuning. However, existing approaches remain inefficient, exhibit limited plasticity, and struggle to accurately preserve unedited regions. To address these issues, we propose MIEdit, a training-free editing framework based on SDE inversion. MIEdit introduces a predictor-corrector multi-history-step scheme to achieve superior editing quality with fewer steps. We further mitigate heterogeneity and conflict between the multi-conditioned noise residuals and gradient terms during sampling, improving stability and editing plasticity under large edits. MIEdit also includes Inversion-Time Automatic Semantic Angle Masking (IASM); it leverages classifier-free guidance to automatically generate semantic angle masks during inversion and applies them throughout the sampling process for regional constraints, without extra user inputs. We additionally construct EditEval++ (30 fine-grained tasks, 1,000+ image-text-mask triplets) for comprehensive evaluation; experiments show that MIEdit outperforms state-of-the-art techniques. Project page: https://whywwwzzzg.github.io/MIEdit/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。