arXiv:2506.07992cs.CV2025-06NeurIPS被引 7

无需文字提示,仅用一对图像即可学习复杂编辑语义。

PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

  • 通过目标噪声预测显式建模图像对间的语义变化方向。
  • 在单对图像下仍能准确学习编辑语义,内容一致性显著提升。
  • 适合需要精细控制但缺乏文本描述的图像编辑场景。

近期基于文本的图像编辑取得了显著进展,但某些语义难以仅凭文本精确描述。现有基于示例的方法仍依赖文本提示或隐式文本指令。本文提出PairEdit,一种无需文本指导、仅需少量图像对(甚至单对)即可有效学习复杂编辑语义的新方法。通过引入目标噪声预测,显式建模图像对间的语义差异,并设计内容保持的噪声调度策略以增强语义学习。同时,采用独立优化的LoRAs分离语义变化与内容信息。大量定性和定量评估表明,PairEdit在保持内容一致性的同时,显著优于基线方法。代码将开源于https://github.com/xudonmao/PairEdit。

原文摘要 · Abstract (English)

Recent advancements in text-guided image editing have achieved notable success by leveraging natural language prompts for fine-grained semantic control. However, certain editing semantics are challenging to specify precisely using textual descriptions alone. A practical alternative involves learning editing semantics from paired source-target examples. Existing exemplar-based editing methods still rely on text prompts describing the change within paired examples or learning implicit text-based editing instructions. In this paper, we introduce PairEdit, a novel visual editing method designed to effectively learn complex editing semantics from a limited number of image pairs or even a single image pair, without using any textual guidance. We propose a target noise prediction that explicitly models semantic variations within paired images through a guidance direction term. Moreover, we introduce a content-preserving noise schedule to facilitate more effective semantic learning. We also propose optimizing distinct LoRAs to disentangle the learning of semantic variations from content. Extensive qualitative and quantitative evaluations demonstrate that PairEdit successfully learns intricate semantics while significantly improving content consistency compared to baseline methods. Code will be available at https://github.com/xudonmao/PairEdit.

图像编辑示例驱动无文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。