评测知识密集型文生图,发现模型常犯科学错误,提出新方法纠正。
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation

- 构建六学科1800个专家级提示,系统评估文生图知识准确性。
- 14款主流模型普遍存在逻辑错误和符号偏差,开源模型表现更差。
- 提出KE-Check框架,通过结构化提示与检查清单提升科学正确性。
近期文生图模型在写实合成与指令遵循方面表现出色,但在知识密集型场景下的可靠性仍待探索。与自然图像生成不同,知识可视化不仅需语义对齐,还要求严格遵守领域知识、结构约束与符号规范,暴露出视觉合理性和科学正确性之间的关键差距。为此,我们提出KVBench——一个基于课程设计的基准,用于评估知识密集型文生图生成。该基准涵盖生物学、化学、地理、历史、数学和物理六门高中课程,包含1800个来自30余本权威教材的专家标注提示。利用此基准,我们评估了14种先进开源与闭源模型,发现其在逻辑推理、符号精确度及多语言鲁棒性方面存在显著缺陷,且开源模型持续落后于闭源系统。为解决这些问题,我们进一步提出KE-Check,一种两阶段框架:(1) 知识扩展,实现提示结构化增强;(2) 检查清单引导精炼,通过违规识别与约束引导编辑实现显式约束强制。KE-Check有效缓解科学幻觉,缩小了开源与领先闭源模型间的性能差距。数据与代码已公开于https://github.com/zhaoran66/KVBench。
原文摘要 · Abstract (English)
Recent text-to-image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliability in knowledge-intensive settings remains largely unexplored. Unlike natural image generation, knowledge visualization requires not only semantic alignment but also strict adherence to domain knowledge, structural constraints, and symbolic conventions, exposing a critical gap between visual plausibility and scientific correctness. To systematically study this problem, we introduce KVBench, a curriculum-grounded benchmark for evaluating knowledge-intensive T2I generation. KVBench covers six senior high-school subjects: Biology, Chemistry, Geography, History, Mathematics, and Physics. The benchmark consists of 1,800 expert-curated prompts derived from over 30 authoritative textbooks. Using this benchmark, we evaluate 14 state-of-the-art open- and closed-source models, revealing substantial deficiencies in logical reasoning, symbolic precision, and multilingual robustness, with open-source models consistently underperforming proprietary systems. To address these limitations, we further propose KE-Check, a two-stage framework that improves scientific fidelity via (1) Knowledge Elaboration for structured prompt enrichment, and (2) Checklist-Guided Refinement for explicit constraint enforcement through violation identification and constraint-guided editing. KE-Check effectively mitigates scientific hallucinations, narrowing the performance gap between open-source and leading closed-source models. Data and codes are publicly available at https://github.com/zhaoran66/KVBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。