arXiv:2503.07265cs.CVcs.AI2025-03中稿 · ICML被引 198

首个评估文生图模型世界知识理解能力的基准,揭示现有模型在常识推理上的短板。

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

  • 设计1000个跨25子领域的高阶语义提示,测试模型对常识、时空和科学知识的理解
  • 提出新指标WiScore,比传统CLIP更精准衡量知识-图像对齐程度
  • 评测20个模型发现其世界知识整合能力普遍不足,适合研究生成质量与认知对齐的学者

文生图(T2I)模型能够生成高质量的艺术创作和视觉内容。然而,现有研究与评估标准主要关注图像真实性和浅层文本-图像对齐,缺乏对复杂语义理解与世界知识融合的全面评估。为解决这一挑战,我们提出首个专用于世界知识感知语义评估的基准——WISE。WISE超越简单的词-像素映射,通过1000个精心设计的提示,覆盖文化常识、时空推理和自然科学等25个子领域。为克服传统CLIP度量的局限,我们引入新的量化指标WiScore,用于评估知识-图像对齐。通过对20个模型(10个专用文生图模型与10个统一多模态模型)使用1000个结构化提示进行综合测试,发现它们在生成过程中有效整合与应用世界知识的能力存在显著局限,揭示了下一代文生图模型提升知识融合与应用的关键路径。代码与数据已公开于PKU-YuanGroup/WISE。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and world knowledge integration in text-to-image generation. To address this challenge, we propose \textbf{WISE}, the first benchmark specifically designed for \textbf{W}orld Knowledge-\textbf{I}nformed \textbf{S}emantic \textbf{E}valuation. WISE moves beyond simple word-pixel mapping by challenging models with 1000 meticulously crafted prompts across 25 subdomains in cultural common sense, spatio-temporal reasoning, and natural science. To overcome the limitations of traditional CLIP metric, we introduce \textbf{WiScore}, a novel quantitative metric for assessing knowledge-image alignment. Through comprehensive testing of 20 models (10 dedicated T2I models and 10 unified multimodal models) using 1,000 structured prompts spanning 25 subdomains, our findings reveal significant limitations in their ability to effectively integrate and apply world knowledge during image generation, highlighting critical pathways for enhancing knowledge incorporation and application in next-generation T2I models. Code and data are available at \href{https://github.com/PKU-YuanGroup/WISE}{PKU-YuanGroup/WISE}.

文生图语义评估世界知识多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。