arXiv:2508.09987cs.CVcs.AI2025-08被引 106

用GPT-4o生成的合成图像提升开源模型,解决真实数据缺失的想象力场景问题。

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

  • 利用GPT-4o生成180K张合成图像,补足真实数据中罕见幻想类场景
  • 合成图像背景纯净、图文对齐精准,提升文本到图像生成质量
  • 适合作为多模态生成模型训练数据,尤其适合需要创意生成的场景

GPT-4o在图像生成方面表现突出,但开源模型仍落后。现有研究尝试从GPT-4o蒸馏图像数据以提升性能,但未解释为何需使用合成数据。本文指出合成图像两大优势:一是补充真实数据中稀缺的奇幻、多参考等复杂场景;二是提供干净可控的监督信号,避免真实数据中的背景噪声与图文错位。基于此,我们构建了180K规模的Echo-4o-Image合成数据集,并以此微调基线模型Bagel,得到Echo-4o。同时提出两个新评测基准:GenEval++通过增强指令复杂度缓解评分饱和,Imagine-Bench专注评估想象力理解与生成能力。实验显示Echo-4o在标准基准上表现优异,且在OmniGen2、BLIP3-o等模型上均带来一致性能提升,证明其强迁移性。

原文摘要 · Abstract (English)

Recently, GPT-4o has garnered significant attention for its strong performance in image generation, yet open-source models still lag behind. Several studies have explored distilling image data from GPT-4o to enhance open-source models, achieving notable progress. However, a key question remains: given that real-world image datasets already constitute a natural source of high-quality data, why should we use GPT-4o-generated synthetic data? In this work, we identify two key advantages of synthetic images. First, they can complement rare scenarios in real-world datasets, such as surreal fantasy or multi-reference image generation, which frequently occur in user queries. Second, they provide clean and controllable supervision. Real-world data often contains complex background noise and inherent misalignment between text descriptions and image content, whereas synthetic images offer pure backgrounds and long-tailed supervision signals, facilitating more accurate text-to-image alignment. Building on these insights, we introduce Echo-4o-Image, a 180K-scale synthetic dataset generated by GPT-4o, harnessing the power of synthetic image data to address blind spots in real-world coverage. Using this dataset, we fine-tune the unified multimodal generation baseline Bagel to obtain Echo-4o. In addition, we propose two new evaluation benchmarks for a more accurate and challenging assessment of image generation capabilities: GenEval++, which increases instruction complexity to mitigate score saturation, and Imagine-Bench, which focuses on evaluating both the understanding and generation of imaginative content. Echo-4o demonstrates strong performance across standard benchmarks. Moreover, applying Echo-4o-Image to other foundation models (e.g., OmniGen2, BLIP3-o) yields consistent performance gains across multiple metrics, highlighting the datasets strong transferability.

图像生成合成数据多模态GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。