构建120万张高质量图像编辑数据集,推动开源模型进步
ImgEdit: A Unified Image Editing Dataset and Benchmark
- 构建多阶段质检流程,确保120万组编辑对的高精度与多样性
- 训练出的ImgEdit-E1模型在多任务上超越现有开源模型
- 提供单轮与多轮挑战性评测套件,适合模型开发者与研究者使用
生成模型在文本到图像生成方面取得显著进展,但开源图像编辑模型仍落后于闭源方案,主要受限于高质量数据和评估基准不足。为此,我们推出ImgEdit,一个包含120万组精心筛选编辑对的大规模高质量图像编辑数据集,涵盖新颖且复杂的单轮编辑及具有挑战性的多轮任务。为保障数据质量,采用融合前沿视觉语言模型、检测模型、分割模型以及特定任务修复流程的多阶段流水线,并辅以严格后处理。ImgEdit在任务新颖性和数据质量上均优于现有数据集。基于此,我们训练了ImgEdit-E1模型,利用视觉语言模型处理参考图像与编辑提示,在多个任务上表现超越现有开源模型,凸显了数据集与模型设计的价值。为进一步评估性能,我们构建ImgEdit-Bench评测基准,涵盖基础测试集、复杂单轮套件和专门多轮套件,对开源与闭源模型及ImgEdit-E1进行了全面评估,揭示当前图像编辑模型的行为特征并提供可操作洞察。数据已公开于https://github.com/PKU-YuanGroup/ImgEdit。
原文摘要 · Abstract (English)
Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks. To overcome these limitations, we introduce ImgEdit, a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks. To ensure the data quality, we employ a multi-stage pipeline that integrates a cutting-edge vision-language model, a detection model, a segmentation model, alongside task-specific in-painting procedures and strict post-processing. ImgEdit surpasses existing datasets in both task novelty and data quality. Using ImgEdit, we train ImgEdit-E1, an editing model using Vision Language Model to process the reference image and editing prompt, which outperforms existing open-source models on multiple tasks, highlighting the value of ImgEdit and model design. For comprehensive evaluation, we introduce ImgEdit-Bench, a benchmark designed to evaluate image editing performance in terms of instruction adherence, editing quality, and detail preservation. It includes a basic testsuite, a challenging single-turn suite, and a dedicated multi-turn suite. We evaluate both open-source and proprietary models, as well as ImgEdit-E1, providing deep analysis and actionable insights into the current behavior of image-editing models. The source data are publicly available on https://github.com/PKU-YuanGroup/ImgEdit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。