arXiv:2505.05501cs.CVcs.AI2025-05被引 4

GPT-4o可直接生成图像,但空间与时间推理仍不靠谱。

Preliminary Explorations with GPT-4o(mni) Native Image Generation

  • 基于多任务分类框架,测试GPT-4o在六类图像生成任务中的表现。
  • 文本到图像生成质量高,但空间定位与时间连贯性严重不足。
  • 适合创意设计,但科学绘图等专业场景易出幻觉与错误。

近期,OpenAI 解锁了 GPT-4o(mni) 的视觉生成能力,其展现出卓越的多模态理解与多样化指令响应能力。本文旨在全面探索 GPT-4o 在各类任务中的表现。受先前研究启发,我们构建了一个任务分类体系,并设计了精心筛选的测试样本,开展定性评估。得益于 GPT-4o 强大的多模态理解能力,其图像生成过程在通用合成任务中表现出超越传统生成模型的能力。我们从模型能力维度出发,在六类任务中评估其性能:传统图像生成、判别类任务、知识驱动生成、常识驱动生成、空间感知生成及时间感知生成。这些任务不仅评估输出质量与条件对齐度,更深入探查模型对现实概念的理解。结果表明,GPT-4o 在文本到图像生成、视觉风格化和低层图像处理方面表现优异;但在精确空间推理、指令对齐生成和一致的时间预测方面存在显著局限。此外,在知识密集或领域特定场景(如科学插图、数学图表)中,模型常出现幻觉、事实错误或结构不一致。这表明,尽管 GPT-4o 在统一多模态生成上实现重大进展,但在专业或安全关键领域仍难以可靠应用。

原文摘要 · Abstract (English)

Recently, the visual generation ability by GPT-4o(mni) has been unlocked by OpenAI. It demonstrates a very remarkable generation capability with excellent multimodal condition understanding and varied task instructions. In this paper, we aim to explore the capabilities of GPT-4o across various tasks. Inspired by previous study, we constructed a task taxonomy along with a carefully curated set of test samples to conduct a comprehensive qualitative test. Benefiting from GPT-4o's powerful multimodal comprehension, its image-generation process demonstrates abilities surpassing those of traditional image-generation tasks. Thus, regarding the dimensions of model capabilities, we evaluate its performance across six task categories: traditional image generation tasks, discriminative tasks, knowledge-based generation, commonsense-based generation, spatially-aware image generation, and temporally-aware image generation. These tasks not only assess the quality and conditional alignment of the model's outputs but also probe deeper into GPT-4o's understanding of real-world concepts. Our results reveal that GPT-4o performs impressively well in general-purpose synthesis tasks, showing strong capabilities in text-to-image generation, visual stylization, and low-level image processing. However, significant limitations remain in its ability to perform precise spatial reasoning, instruction-grounded generation, and consistent temporal prediction. Furthermore, when faced with knowledge-intensive or domain-specific scenarios, such as scientific illustrations or mathematical plots, the model often exhibits hallucinations, factual errors, or structural inconsistencies. These findings suggest that while GPT-4o marks a substantial advancement in unified multimodal generation, there is still a long way to go before it can be reliably applied to professional or safety-critical domains.

图像生成多模态GPT-4o评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。