arXiv:2506.10963cs.CVcs.CL2025-06NeurIPS被引 14

评测文生图模型的知识推理能力,提出新基准MMMG

MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning

  • 构建跨10学科、6教育层级的图文生成评测集
  • 用知识图谱量化图像事实准确性和视觉清晰度
  • 发现主流模型存在实体缺失、关系薄弱等缺陷

本文提出知识图像生成新任务,并发布大规模多学科多层级知识-图像生成基准MMMG,用于评估图像生成模型的推理能力。知识图像在人类文明与学习中至关重要,其生成需融合世界知识与像素级对齐。MMMG包含4,456组专家验证的图文对,覆盖10个学科、6个教育层次及图表、思维导图等多种知识形式。为消除评估干扰,采用统一知识图谱(KG)表示,明确目标图像的核心实体及其依赖关系。引入MMMG-Score评估生成图像质量,结合图编辑距离衡量事实准确性与视觉清晰度评估。对16个先进文生图模型的全面测试显示,模型普遍存在实体保真度低、关系弱、画面杂乱等问题,GPT-4o仅获50.20分。为推动研究,我们发布开源基线FLUX-Reason,结合推理大模型与扩散模型,在16,000组精选数据上训练,获34.45分。

原文摘要 · Abstract (English)

In this paper, we introduce knowledge image generation as a new task, alongside the Massive Multi-Discipline Multi-Tier Knowledge-Image Generation Benchmark (MMMG) to probe the reasoning capability of image generation models. Knowledge images have been central to human civilization and to the mechanisms of human learning -- a fact underscored by dual-coding theory and the picture-superiority effect. Generating such images is challenging, demanding multimodal reasoning that fuses world knowledge with pixel-level grounding into clear explanatory visuals. To enable comprehensive evaluation, MMMG offers 4,456 expert-validated (knowledge) image-prompt pairs spanning 10 disciplines, 6 educational levels, and diverse knowledge formats such as charts, diagrams, and mind maps. To eliminate confounding complexity during evaluation, we adopt a unified Knowledge Graph (KG) representation. Each KG explicitly delineates a target image's core entities and their dependencies. We further introduce MMMG-Score to evaluate generated knowledge images. This metric combines factual fidelity, measured by graph-edit distance between KGs, with visual clarity assessment. Comprehensive evaluations of 16 state-of-the-art text-to-image generation models expose serious reasoning deficits -- low entity fidelity, weak relations, and clutter -- with GPT-4o achieving an MMMG-Score of only 50.20, underscoring the benchmark's difficulty. To spur further progress, we release FLUX-Reason (MMMG-Score of 34.45), an effective and open baseline that combines a reasoning LLM with diffusion models and is trained on 16,000 curated knowledge image-prompt pairs.

文生图知识推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。