构建40万张真实图像的文本编辑数据集,推动下一代图像编辑模型发展。
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
- 用Nano-Banana生成真实照片的多样化编辑对,确保内容忠实与指令精准。
- 包含72万次多轮编辑、56万条偏好数据,支持复杂编辑研究。
- 适合研究图像编辑、指令对齐与模型评估的开发者和研究人员。
近期多模态模型在文本引导图像编辑方面取得显著进展,如GPT-4o和Nano-Banana已达到新基准。然而,研究进展受限于缺乏大规模、高质量且公开可用的真实图像数据集。本文提出Pico-Banana-400K,一个包含40万张图像的指令式图像编辑数据集。该数据集基于OpenImages中的真实照片,利用Nano-Banana生成多样化的编辑对。与以往合成数据集不同,本研究通过细粒度编辑分类体系确保覆盖全面的编辑类型,并采用基于多模态大模型(MLLM)的质量评分与人工精炼,保障内容保留精度与指令一致性。数据集还包含三个专项子集:(1) 7.2万样本的多轮编辑集,用于研究连续修改中的推理与规划;(2) 5.6万样本的偏好数据集,支持对齐研究与奖励模型训练;(3) 配对的长/短指令对,用于开发指令重写与摘要能力。Pico-Banana-400K为训练和评测下一代文本引导图像编辑模型提供了高质量、任务丰富的基础资源。
原文摘要 · Abstract (English)
Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset is constructed by leveraging Nano-Banana to generate diverse edit pairs from real photographs in the OpenImages collection. What distinguishes Pico-Banana-400K from previous synthetic datasets is our systematic approach to quality and diversity. We employ a fine-grained image editing taxonomy to ensure comprehensive coverage of edit types while maintaining precise content preservation and instruction faithfulness through MLLM-based quality scoring and careful curation. Beyond single turn editing, Pico-Banana-400K enables research into complex editing scenarios. The dataset includes three specialized subsets: (1) a 72K-example multi-turn collection for studying sequential editing, reasoning, and planning across consecutive modifications; (2) a 56K-example preference subset for alignment research and reward model training; and (3) paired long-short editing instructions for developing instruction rewriting and summarization capabilities. By providing this large-scale, high-quality, and task-rich resource, Pico-Banana-400K establishes a robust foundation for training and benchmarking the next generation of text-guided image editing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。