arXiv:2507.21033cs.CV2025-07被引 69

用GPT-4o生成150万张高质量图像编辑数据,推动开源研究

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

  • 用GPT-4o统一优化三个编辑数据集,提升图像质量和指令对齐
  • 在多个基准上实现7.24至8.78的高分,超越现有开源方法
  • 适合做图像编辑、指令跟随等方向的研究者使用

近年来,GPT-4o等大型多模态模型在高保真、指令引导的图像编辑方面树立了新标准。然而,这些模型及其训练数据的专有性给开源研究带来了显著障碍。为弥合这一差距,我们提出GPT-IMAGE-EDIT-1.5M,一个公开可用的大规模图像编辑语料库,包含超过150万条高质量三元组(指令、源图像、编辑后图像)。我们通过利用GPT-4o的多功能性,系统性地整合并优化了三个主流图像编辑数据集:OmniEdit、HQ-Edit和UltraEdit。具体方法包括:1)重新生成输出图像以提升视觉质量与指令一致性;2)选择性重写提示词以增强语义清晰度。为验证数据集的有效性,我们在GPT-IMAGE-EDIT-1.5M上微调先进的开源模型。实验结果令人振奋,例如,微调后的FluxKontext在多项基准测试中表现优异,包括在GEdit-EN上达到7.24,在ImgEdit-Full上为3.80,在Complex-Edit上为8.78,展现出更强的指令遵循能力与更高的感知质量,同时保持身份一致性。这些分数显著超越此前所有已发布的开源方法,并大幅缩小与领先专有模型之间的差距。我们希望完整发布GPT-IMAGE-EDIT-1.5M能进一步推动指令引导图像编辑的开源研究。

原文摘要 · Abstract (English)

Recent advancements in large multimodal models like GPT-4o have set a new standard for high-fidelity, instruction-guided image editing. However, the proprietary nature of these models and their training data creates a significant barrier for open-source research. To bridge this gap, we introduce GPT-IMAGE-EDIT-1.5M, a publicly available, large-scale image-editing corpus containing more than 1.5 million high-quality triplets (instruction, source image, edited image). We systematically construct this dataset by leveraging the versatile capabilities of GPT-4o to unify and refine three popular image-editing datasets: OmniEdit, HQ-Edit, and UltraEdit. Specifically, our methodology involves 1) regenerating output images to enhance visual quality and instruction alignment, and 2) selectively rewriting prompts to improve semantic clarity. To validate the efficacy of our dataset, we fine-tune advanced open-source models on GPT-IMAGE-EDIT-1.5M. The empirical results are exciting, e.g., the fine-tuned FluxKontext achieves highly competitive performance across a comprehensive suite of benchmarks, including 7.24 on GEdit-EN, 3.80 on ImgEdit-Full, and 8.78 on Complex-Edit, showing stronger instruction following and higher perceptual quality while maintaining identity. These scores markedly exceed all previously published open-source methods and substantially narrow the gap to leading proprietary models. We hope the full release of GPT-IMAGE-EDIT-1.5M can help to catalyze further open research in instruction-guided image editing.

图像编辑指令跟随数据集GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。