GPT-4o生成图像时常忽略深层语义,缺乏常识推理能力。
Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
- 通过三维度评估模型对指令、细节编辑和推理任务的响应。
- 80%以上指令理解偏差,知识约束应用不一致,条件推理失败率超60%。
- 适合研究多模态生成与认知对齐的学者参考。
OpenAI的多模态模型GPT-4o在图像生成与编辑方面表现出色,但其是否具备融合世界知识、上下文推理与指令遵循的语义合成能力仍待验证。本研究从三个关键维度系统评估:(1)全局指令遵循性,(2)细粒度编辑精度,(3)生成后推理能力。尽管现有基准显示其生成能力强,我们的评估揭示其存在持续局限:模型常采用字面解释指令,知识约束应用不一致,且在条件推理任务中表现不佳。这些发现挑战了关于GPT-4o统一理解与生成能力的普遍认知,暴露出其动态知识整合的重大缺口。研究呼吁开发更鲁棒的基准与训练策略,超越表面对齐,强调上下文感知与推理驱动的多模态生成。
原文摘要 · Abstract (English)
OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contextual reasoning, and instruction adherence--remains unproven. In this study, we systematically evaluate these capabilities across three critical dimensions: (1) Global Instruction Adherence, (2) Fine-Grained Editing Precision, and (3) Post-Generation Reasoning. While existing benchmarks highlight GPT-4o's strong capabilities in image generation and editing, our evaluation reveals GPT-4o's persistent limitations: the model frequently defaults to literal interpretations of instructions, inconsistently applies knowledge constraints, and struggles with conditional reasoning tasks. These findings challenge prevailing assumptions about GPT-4o's unified understanding and generation capabilities, exposing significant gaps in its dynamic knowledge integration. Our study calls for the development of more robust benchmarks and training strategies that go beyond surface-level alignment, emphasizing context-aware and reasoning-grounded multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。