构建大规模图像生成编辑数据集,提升模型真实场景表现
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
- 按层级任务体系自动构建8万条高质量指令-图像对
- 在编辑与生成任务上分别提升18%和13%性能
- 适合研究多模态生成、编辑与数据构建的学者
统一多模态模型在图像生成与编辑中的表现受限于训练数据的质量与全面性。现有数据集虽涵盖风格迁移、简单物体操作等基础任务,但缺乏系统结构与现实挑战场景。为此,我们提出OpenGPT-4o-Image,一种基于层次化任务分类与自动化生成的新方法构建的大规模数据集。该分类不仅包含文本渲染、风格控制等基础能力,还引入化学图示等实用且具挑战性的类别,以及需同时执行多项操作的复杂指令编辑。通过结合结构化资源池与GPT-4o的自动化流水线,生成8万条高质指令-图像对,覆盖11个主要领域与51个子任务。大量实验表明,在该数据集上微调主流模型可显著提升多个基准的表现,编辑任务(UniWorld-V1 on ImgEdit-Bench)最高提升18%,生成任务(Harmon on GenEval)提升13%。结果表明,系统性数据构建是推动多模态AI能力发展的关键。
原文摘要 · Abstract (English)
The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18\% on editing tasks (UniWorld-V1 on ImgEdit-Bench) and 13% on generation tasks (Harmon on GenEval). Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。