arXiv:2506.18095cs.CVcs.AI2025-06被引 97

用GPT-4o生成数据训练出开源图像模型Janus-4o

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

  • 用GPT-4o合成45K图文数据,蒸馏其生成能力
  • 仅用91K样本训练6小时即实现高质量图文生成
  • 首个支持文本+图像生成新图像的开源多模态模型

多模态生成模型虽已实现逼真、指令对齐的图像生成,但如GPT-4o-Image等先进系统仍为私有不可用。为推动开放研究,我们提出ShareGPT-4o-Image,首个包含45,000条文生图与46,000条图文生图数据的开源数据集,均通过GPT-4o生成。基于该数据集,我们开发了Janus-4o,一个兼具文生图与图文生图能力的多模态大模型。相比前代Janus-Pro,其文生图性能显著提升,并首次实现图文生图功能。令人惊喜的是,仅用91,000条合成样本和8张A800 GPU上6小时训练,即可达到优异表现。我们希望分享该数据集与模型,促进逼真、指令对齐图像生成的开放研究。

原文摘要 · Abstract (English)

Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabilities, we present ShareGPT-4o-Image, the first dataset comprising 45K text-to-image and 46K text-and-image-to-image data, all synthesized using GPT-4o's image generation capabilities for distilling its advanced image generation abilities. Leveraging this dataset, we develop Janus-4o, a multimodal large language model capable of both text-to-image and text-and-image-to-image generation. Janus-4o not only significantly improves text-to-image generation over its predecessor, Janus-Pro, but also newly supports text-and-image-to-image generation. Notably, it achieves impressive performance in text-and-image-to-image generation from scratch, using only 91K synthetic samples and 6 hours of training on an 8 A800-GPU machine. We hope the release of ShareGPT-4o-Image and Janus-4o will foster open research in photorealistic, instruction-aligned image generation.

多模态生成图文生成开源模型图像对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。