首个评估生成模型隐性世界知识的基准,揭示现有模型在物理推理上的普遍短板。
Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- 构建跨三类场景的1100个提示,测试隐性知识与物理因果推理能力。
- 提出多智能体评估框架PW-Agent,逐层验证图像的物理真实性和逻辑一致性。
- 17个主流文生图模型均暴露推理缺陷,亟需融合知识与因果的新型架构。
当前文本到图像(T2I)模型虽能生成逼真且遵循指令的图像,但在涉及隐性世界知识的提示上仍频繁失败。现有评估协议或侧重组合对齐,或依赖单轮VQA评分,对知识定位、多物理交互及可审计证据等关键维度严重不足。为此,我们提出PicWorld——首个全面评估T2I模型隐性世界知识与物理因果推理能力的基准,包含1100个跨三大类别的提示。为支持细粒度评估,我们设计了PW-Agent:一个基于证据的多智能体评估器,通过分解提示为可验证视觉证据,分层判断图像的物理真实性与逻辑一致性。对17个主流T2I模型的全面分析表明,它们在隐性知识与物理因果推理方面普遍存在根本性局限。研究强调未来T2I系统亟需引入推理感知与知识整合架构。
原文摘要 · Abstract (English)
Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize compositional alignment or rely on single-round VQA-based scoring, leaving critical dimensions such as knowledge grounding, multi-physics interactions, and auditable evidence-substantially undertested. To address these limitations, we introduce PicWorld, the first comprehensive benchmark that assesses the grasp of implicit world knowledge and physical causal reasoning of T2I models. This benchmark consists of 1,100 prompts across three core categories. To facilitate fine-grained evaluation, we propose PW-Agent, an evidence-grounded multi-agent evaluator to hierarchically assess images on their physical realism and logical consistency by decomposing prompts into verifiable visual evidence. We conduct a thorough analysis of 17 mainstream T2I models on PicWorld, illustrating that they universally exhibit a fundamental limitation in their capacity for implicit world knowledge and physical causal reasoning to varying degrees. The findings highlight the need for reasoning-aware, knowledge-integrative architectures in future T2I systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。