arXiv:2509.14232cs.CV2025-09中稿 · ICML被引 14

首个跨学科图文生成考试基准,评估模型综合理解与生成能力。

GenExam: A Multidisciplinary Text-to-Image Exam

  • 构建10个学科、1000题的考试式图文生成数据集,含四层分类体系。
  • 每题配真实图像与细粒度评分点,精准衡量语义正确性与视觉合理性。
  • 开源模型在该测试中显著落后于闭源领先模型,揭示生成瓶颈。

考试是检验专家级智能的核心方式,需整合理解、推理与生成能力。现有考试类评测多聚焦理解与推理,而生成类基准则偏重世界知识与视觉概念呈现,忽视对严谨绘图能力的评估。我们提出GenExam,首个跨学科文本到图像考试基准,包含10个学科、1000个样本,采用四层分类体系组织考试式提示。每个题目配有真值图像与细粒度评分点,支持对语义准确性和视觉合理性的精确评估。在17个文本到图像及统一模型上的实验表明,GenExam具有极大挑战性,开源模型始终显著落后于领先闭源模型。通过将图像生成置于考试情境下,GenExam为评估模型融合理解、推理与生成的能力提供了严格标准,推动智能生成模型的发展。相关基准与评估代码已开源:https://github.com/OpenGVLab/GenExam。

原文摘要 · Abstract (English)

Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style benchmarks mainly focus on understanding and reasoning tasks, and current generation benchmarks emphasize the illustration of world knowledge and visual concepts, neglecting the evaluation of rigorous drawing exams. We introduce GenExam, the first benchmark for multidisciplinary text-to-image exams, featuring 1,000 samples across 10 subjects with exam-style prompts organized under a four-level taxonomy. Each problem is equipped with ground-truth images and fine-grained scoring points to enable a precise evaluation of semantic correctness and visual plausibility. Experiments on 17 text-to-image and unified models demonstrate the great challenge of GenExam and the huge gap where open-source models consistently lag behind the leading closed-source ones. By framing image generation as an exam, GenExam offers a rigorous assessment of models' ability to integrate understanding, reasoning, and generation, providing insights for on the path to intelligent generative models. Our benchmark and evaluation code are released at https://github.com/OpenGVLab/GenExam.

图文生成考试评测多学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。