通过双条件引导提升扩散模型图像编辑的精度与灵活性
DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing
- 联合源提示与参考图指导反演过程,稳定语义与视觉空间轨迹
- 在多类编辑任务中实现最优重建质量与编辑精度
- 适合需要高保真图像编辑与鲁棒反演的应用场景
扩散模型在图像生成与编辑任务中取得显著进展。图像反演旨在恢复真实或生成图像的潜在噪声表示,以支持重建、编辑等下游任务。然而,现有反演方法普遍存在重建精度与编辑灵活性之间的固有权衡,这源于反演过程中保持语义对齐与结构一致性之难。本文提出双条件反演(DCI),通过联合源提示与参考图像引导反演过程,将反演建模为双重条件下的固定点优化问题,同时最小化潜在噪声差距与重构误差。该设计在语义与视觉空间双重锚定反演路径,获得更准确且可编辑的潜在表示。大量实验表明,DCI在多项编辑任务中达到当前最佳性能,显著提升重建质量与编辑精度。此外,该方法在重建任务中也表现优异,展现出接近反演终极目标的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable success in image generation and editing tasks. Inversion within these models aims to recover the latent noise representation for a real or generated image, enabling reconstruction, editing, and other downstream tasks. However, to date, most inversion approaches suffer from an intrinsic trade-off between reconstruction accuracy and editing flexibility. This limitation arises from the difficulty of maintaining both semantic alignment and structural consistency during the inversion process. In this work, we introduce Dual-Conditional Inversion (DCI), a novel framework that jointly conditions on the source prompt and reference image to guide the inversion process. Specifically, DCI formulates the inversion process as a dual-condition fixed-point optimization problem, minimizing both the latent noise gap and the reconstruction error under the joint guidance. This design anchors the inversion trajectory in both semantic and visual space, leading to more accurate and editable latent representations. Our novel setup brings new understanding to the inversion process. Extensive experiments demonstrate that DCI achieves state-of-the-art performance across multiple editing tasks, significantly improving both reconstruction quality and editing precision. Furthermore, we also demonstrate that our method achieves strong results in reconstruction tasks, implying a degree of robustness and generalizability approaching the ultimate goal of the inversion process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。