arXiv:2411.17323cs.CV2024-11CVPR被引 21

提升图像编辑对复杂指令的理解能力,保持背景一致性。

InsightEdit: Towards Better Instruction Following for Image Editing

  • 构建新数据集AdvancedEdit,解决原数据低分辨率与指令简单问题。
  • 引入双流桥接机制,融合文本与视觉特征,精准指导编辑过程。
  • 适合需要高保真图像编辑的AI研究者与开发者使用。

本文聚焦于基于指令的图像编辑任务。以往工作如InstructPix2Pix、InstructDiffusion和SmartEdit虽已探索端到端编辑,但仍存在两大局限:一是现有数据集分辨率低、背景一致性差、指令过于简单;二是当前方法主要依赖文本条件,忽视丰富的图像信息,导致在复杂指令理解与背景一致性维持方面表现不佳。针对上述问题,我们首先通过新颖的数据构建流程,构建了AdvancedEdit数据集,该数据集具备高视觉质量、复杂指令与良好的背景一致性。随后,为更充分挖掘图像信息,提出一种基于多模态大模型(MLLM)推理出的文本与视觉特征的双流桥接机制,以更精确地引导图像编辑过程。大量实验表明,所提方法InsightEdit达到当前最优性能,在复杂指令遵循与保持原始图像背景一致性方面表现优异。

原文摘要 · Abstract (English)

In this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, InstructDiffusion, and SmartEdit have explored end-to-end editing. However, two limitations still remain: First, existing datasets suffer from low resolution, poor background consistency, and overly simplistic instructions. Second, current approaches mainly condition on the text while the rich image information is underexplored, therefore inferior in complex instruction following and maintaining background consistency. Targeting these issues, we first curated the AdvancedEdit dataset using a novel data construction pipeline, formulating a large-scale dataset with high visual quality, complex instructions, and good background consistency. Then, to further inject the rich image information, we introduce a two-stream bridging mechanism utilizing both the textual and visual features reasoned by the powerful Multimodal Large Language Models (MLLM) to guide the image editing process more precisely. Extensive results demonstrate that our approach, InsightEdit, achieves state-of-the-art performance, excelling in complex instruction following and maintaining high background consistency with the original image.

图像编辑多模态指令跟随生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。