arXiv:2606.16457cs.CVcs.GR2026-06中稿 · the EGSR 2026 jour…

通过残差嵌入提升图像编辑精度与身份一致性。

ResEdit: Residual embeddings for precise generative image editing

论文配图:ResEdit: Residual embeddings for precise generative image editing
图 1 · 摘自论文原文
  • 引入残差图像编码作为额外条件,增强重建信号。
  • 实现高保真内在编辑与光照调整,保持全局一致性。
  • 适合需要精准控制的图像编辑任务,如风格迁移与重光照。

条件扩散图像生成器可通过反演技术用于图像编辑,无需大规模成对微调数据。然而,在保持图像身份和全局一致性的同时实现高质量、精准编辑仍具挑战,因弱条件反演常将冲突特征嵌入噪声中。本文证明,引入残差图像编码作为附加条件,可同时改善身份保留与编辑能力。通过优化该残差编码以提供强重建信号,降低对反演的依赖及其固有缺陷。为避免残差干扰编辑目标,采用基于梯度反转的优化策略,实现残差与编辑条件解耦。实验表明,该方法在精确的内在编辑与重光照任务中均能生成高保真结果,并验证了文本引导操作的可行性。

原文摘要 · Abstract (English)

Conditional diffusion image generators can be repurposed for editing through inversion, without the need for large-scale paired fine-tuning data. However, producing high-quality, targeted edits while maintaining image identity and global consistency remains challenging, as weakly conditioned inversion often embeds conflicting image features into the noise. We demonstrate that incorporating a residual image encoding as additional conditioning enables both improved identity preservation and better editability. We optimize this residual encoding to provide a strong conditioning signal for reconstruction, thereby reducing the reliance on inversion and susceptibility to its aforementioned pitfalls. To ensure this residual does not interfere with desired edits, we incorporate a gradient reversal-based optimization strategy that disentangles the residual from the edited condition. We illustrate our method's ability to produce high-fidelity results across precise intrinsic-based editing and relighting, and show proof-of-concept text-guided manipulation.

图像编辑扩散模型残差编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。