arXiv:2504.05979cs.CV2025-04被引 29

实证分析GPT-4o图像生成能力,对比主流模型表现。

An Empirical Study of GPT-4o Image Generation Capabilities

  • 通过20多个任务实测GPT-4o多模态生成性能。
  • 在文本到图像、图像到3D等任务中展现高保真生成能力。
  • 揭示统一架构设计对生成模型的关键影响,适合研究者参考。

图像生成领域已从早期的GAN方法发展到扩散模型,最近则趋向于统一的生成架构,实现理解与生成任务的融合。尽管GPT-4o等新模型展现出高保真多模态生成能力,其架构设计仍未公开。本文通过实证研究评估GPT-4o的图像生成能力,对比领先开源与商业模型。评估涵盖文本到图像、图像到图像、图像到3D及图像到X生成四大类,共20余项任务。分析揭示了GPT-4o在不同场景下的优势与局限,为其在生成模型演进中的位置提供定位,并指明未来统一生成模型的发展方向,强调架构设计与数据规模的重要性。

原文摘要 · Abstract (English)

The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to bridge understanding and generation tasks. Recent advances, especially the GPT-4o, have demonstrated the feasibility of high-fidelity multimodal generation, their architectural design remains mysterious and unpublished. This prompts the question of whether image and text generation have already been successfully integrated into a unified framework for those methods. In this work, we conduct an empirical study of GPT-4o's image generation capabilities, benchmarking it against leading open-source and commercial models. Our evaluation covers four main categories, including text-to-image, image-to-image, image-to-3D, and image-to-X generation, with more than 20 tasks. Our analysis highlights the strengths and limitations of GPT-4o under various settings, and situates it within the broader evolution of generative modeling. Through this investigation, we identify promising directions for future unified generative models, emphasizing the role of architectural design and data scaling. For a high-definition version of the PDF, please refer to the link on GitHub: \href{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}.

图像生成多模态GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。