让图像编辑保持结构不变,同时精准遵循文本指令。
Consistent-Inversion: Reverse Consistency Guidance for Structure-Preserving Visual Editing
- 通过反向一致性检测修正编辑轨迹,避免结构失真。
- 在PIE-Bench上提升背景和布局保真度,且不损失指令匹配性。
- 无需训练,兼容主流编辑框架,推理开销小。
文本引导的扩散模型已成为真实图像编辑的有效工具,要求编辑后的图像既符合目标指令,又保留无关编辑的结构信息。现有免训练编辑方法依赖于反演:将源图像映射到噪声潜空间轨迹,并复用终点潜变量进行目标提示去噪。这种复用虽有助于结构保留,却耦合了源图重建与目标编辑,导致轨迹不匹配,可能破坏背景/布局细节或过度约束编辑意图。本文提出Consistent-Inversion,一种免训练的反向一致性引导框架,用于结构保持的视觉编辑。不同于将反演源潜变量视为固定初始化,Consistent-Inversion检查在源提示下,中间目标轨迹能否被逆向回源反演轨迹。为使该检查明确,我们构建辅助的目标侧噪声表示,执行源引导的反向去噪,并以反向一致性偏差作为早期目标去噪步骤的校正信号。该方法不更新模型参数,兼容基于反演的编辑器,稀疏应用时仅引入微小推理开销。在PIE-Bench上的实验表明,Consistent-Inversion在统一的SD3.5协议下提升了背景和结构保真度,同时保持目标提示对齐;兼容性实验进一步验证了该校正原理在经典Stable-Diffusion反演流程中的有效性。
原文摘要 · Abstract (English)
Text-guided diffusion models have become effective tools for real-image visual editing, where the edited image must follow a target instruction while preserving editing-irrelevant structure. Most training-free editors rely on inversion: a source image is mapped to a noisy latent trajectory and the terminal latent is reused for target-prompt denoising. This reuse is useful for preservation, but it also couples source reconstruction and target editing. The resulting trajectory mismatch may either damage background/layout details or over-constrain the intended edit. This paper presents Consistent-Inversion, a training-free reverse consistency guidance framework for structure-preserving visual editing. Instead of treating the inverted source latent as a fixed initialization, Consistent-Inversion checks whether an intermediate target trajectory can be reversed toward the source inversion trajectory under the source prompt. To make this check well-defined, we construct an auxiliary target-side noise representation, perform source-guided reverse denoising, and use the resulting reverse consistency discrepancy as a correction signal for selected early target denoising steps. The method does not update model parameters, is compatible with inversion-based editors, and introduces only a small inference overhead when applied sparsely. Experiments on PIE-Bench show that Consistent-Inversion improves background and structural fidelity under a unified SD3.5 protocol while maintaining target-prompt alignment, and compatibility experiments further verify the same correction principle on classical Stable-Diffusion inversion pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。