arXiv:2412.15216cs.CV2024-12ICCV被引 1

无需真实编辑图像即可训练,实现更精准的指令化图像编辑。

UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint

  • 通过正反编辑一致性约束,实现无监督训练。
  • 在多种编辑任务中表现更优,保持高保真与精确度。
  • 适合追求低成本、高泛化能力的图像编辑研究者。

我们提出一种无监督的指令式图像编辑方法,无需在训练中使用真实编辑后的图像。现有方法依赖于包含输入图像、真实编辑图像和编辑指令的三元组,这些三元组通常由已有编辑方法生成(引入偏差),或通过人工标注(成本高且限制泛化性)。我们的方法引入一种新机制——编辑可逆性约束(Edit Reversibility Constraint, ERC),在单个训练步骤中同时执行正向和反向编辑,并在图像、文本和注意力空间中强制对齐。这使我们无需真实编辑图像,首次可在仅含真实图像-标题对或图像-标题-指令三元组的数据集上进行训练。实验证明,该方法在更广泛的编辑任务中表现更优,具有高保真度与精确度。通过消除对预存三元组数据集的需求,减少现有方法的偏差,并提出ERC,本工作在突破指令式图像编辑的规模化瓶颈方面具有显著进展。

原文摘要 · Abstract (English)

We propose an unsupervised instruction-based image editing approach that removes the need for ground-truth edited images during training. Existing methods rely on supervised learning with triplets of input images, ground-truth edited images, and edit instructions. These triplets are typically generated either by existing editing methods, introducing biases, or through human annotations, which are costly and limit generalization. Our approach addresses these challenges by introducing a novel editing mechanism called Edit Reversibility Constraint (ERC), which applies forward and reverse edits in one training step and enforces alignment in image, text, and attention spaces. This allows us to bypass the need for ground-truth edited images and unlock training for the first time on datasets comprising either real image-caption pairs or image-caption-instruction triplets. We empirically show that our approach performs better across a broader range of edits with high-fidelity and precision. By eliminating the need for pre-existing datasets of triplets, reducing biases associated with current methods, and proposing ERC, our work represents a significant advancement in unblocking scaling of instruction-based image editing.

图像编辑无监督学习指令控制可逆性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。