测试文生图模型对世界知识和隐含推理的掌握能力
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
- 设计新基准评估图像生成中的世界知识与推理能力
- 21个先进模型中扩散模型表现最优,闭源模型更强
- 适合研究图文理解、认知对齐与真实场景生成的学者
文本到图像(T2I)生成虽取得显著进展,但在涉及丰富世界知识和隐含推理的提示上仍表现不佳,而这对于生成语义准确、连贯且符合上下文的真实场景图像至关重要。为此,我们提出 extbf{WorldGenBench},一个系统评估 T2I 模型世界知识锚定与隐含推断能力的基准,涵盖人文与自然领域。我们引入 extbf{Knowledge Checklist Score},一种结构化指标,衡量生成图像是否满足关键语义预期。在 21 个前沿模型上的实验表明,尽管开源扩散模型领先,但专有自回归模型如 GPT-4o 在推理与知识整合方面表现显著更优。研究凸显了下一代 T2I 系统亟需更强的理解与推理能力。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models still struggle with prompts that require rich world knowledge and implicit reasoning: both of which are critical for producing semantically accurate, coherent, and contextually appropriate images in real-world scenarios. To address this gap, we introduce \textbf{WorldGenBench}, a benchmark designed to systematically evaluate T2I models' world knowledge grounding and implicit inferential capabilities, covering both the humanities and nature domains. We propose the \textbf{Knowledge Checklist Score}, a structured metric that measures how well generated images satisfy key semantic expectations. Experiments across 21 state-of-the-art models reveal that while diffusion models lead among open-source methods, proprietary auto-regressive models like GPT-4o exhibit significantly stronger reasoning and knowledge integration. Our findings highlight the need for deeper understanding and inference capabilities in next-generation T2I systems. Project Page: \href{https://dwanzhang-ai.github.io/WorldGenBench/}{https://dwanzhang-ai.github.io/WorldGenBench/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。