arXiv:2505.19352cs.CV2025-05被引 7

用海量图文对训练细粒度图像编辑,无需人工标注编辑对。

Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions

  • 引入多尺度可学习区域定位编辑位置,基于图文对自监督训练。
  • 在多个基准上达到顶尖效果,支持多种生成模型适配。
  • 突破传统依赖人工编辑对的局限,适合通用图像编辑场景。

当前文本驱动的图像编辑方法主要分为两类:依赖大规模高质量编辑对数据集以提升精度与多样性,或探索无需数据集的替代方案。然而,构建大规模编辑数据集需复杂流程、耗时长,常产生不真实样本或伪影;而无数据集方法则存在指令理解能力弱、编辑能力受限的问题。为此,本文提出一种新范式,利用广泛存在的海量文本-图像对,而非依赖专门编辑对数据集。通过将图像与文本描述间的对齐作为监督信号,学习生成特定任务的可编辑区域,实现高保真、精准且符合指令的图像编辑。大量实验表明,该方法在多种任务和基准上均达到最先进性能,且对不同生成模型具有强适应性。

原文摘要 · Abstract (English)

Current text-driven image editing methods typically follow one of two directions: relying on large-scale, high-quality editing pair datasets to improve editing precision and diversity, or exploring alternative dataset-free techniques. However, constructing large-scale editing datasets requires carefully designed pipelines, is time-consuming, and often results in unrealistic samples or unwanted artifacts. Meanwhile, dataset-free methods may suffer from limited instruction comprehension and restricted editing capabilities. Faced with these challenges, the present work develops a novel paradigm for instruction-driven image editing that leverages widely available and enormous text-image pairs, instead of relying on editing pair datasets. Our approach introduces a multi-scale learnable region to localize and guide the editing process. By treating the alignment between images and their textual descriptions as supervision and learning to generate task-specific editing regions, our method achieves high-fidelity, precise, and instruction-consistent image editing. Extensive experiments demonstrate that the proposed approach attains state-of-the-art performance across various tasks and benchmarks, while exhibiting strong adaptability to various types of generative models.

图像编辑文本引导自监督生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。