arXiv:2608.16812cs.CV2026-08

通过细粒度概念与密集监督提升图像编辑精度与效率

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

论文配图:Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
图 1 · 摘自论文原文
  • 构建超1000个细粒度编辑概念的分类体系
  • 训练效率提升,生成质量显著优于现有方法
  • 适合需要高精度图像编辑的研究者与开发者

现有图像编辑框架多沿用文本到图像扩散模型的训练范式,但将其扩展至图像编辑时存在两大问题:编辑概念粒度不足,以及稀疏监督信号导致的训练低效。为此,我们建立了一个包含1000多个细粒度编辑概念的完整层次化分类体系,并通过改进的合成框架构建了规模达1200万对的高质量编辑数据集ConceptEdit-12M。该基于知识库的方法有效缓解了生成数据分布坍塌问题,同时保障高数据保真度。此外,提出一种密集监督训练策略,将多个无冲突概念融合进单一图像对中,提供更丰富的学习信号,显著提升训练效率与模型性能。实验验证表明,该方法显著优于先前工作。最后,我们发布了ConceptEdit-Bench,一个用于诊断模型在多样化真实场景下能力的细粒度评估基准。

原文摘要 · Abstract (English)

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.

图像编辑扩散模型数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。