评测视觉语言模型生成科研核心图的能力,推动多模态AI理解与表达。
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
- 基于论文标题摘要等文本生成科学概念图,需融合理解与创意设计。
- 顶尖模型在生成准确且有说服力的科研图上仍表现不佳。
- 适合关注多模态生成、科学可视化与AI评估的研究者。
许多科学论文中的“图1”是核心研究思想的主要视觉摘要,虽形式简洁但概念丰富,常需作者反复推敲,凸显科学可视化难度。基于此,我们提出GENFIG1基准,用于评估生成式AI模型(如视觉语言模型)生成能清晰表达并支撑论文核心思想的图表能力。输入包括论文标题、摘要、引言和图注。解决GENFIG1不仅要求图像美观,还需结合科学理解进行图文推理:(i)理解论文技术概念,(ii)识别最显著信息,(iii)设计逻辑一致且视觉有效的图形。我们从顶级深度学习会议论文中构建数据集,实施严格质量控制,并引入一个与专家判断高度相关的自动评估指标。对多种代表性模型的评估显示,该任务对当前最佳系统仍是重大挑战。我们希望这一基准能为多模态AI未来发展奠定基础。
原文摘要 · Abstract (English)
In many science papers, "Figure 1" serves as the primary visual summary of the core research idea. These figures are visually simple yet conceptually rich, often requiring significant effort and iteration by human authors to get right, highlighting the difficulty of science visual communication. With this intuition, we introduce GENFIG1, a benchmark for generative AI models (e.g., Vision-Language Models). GENFIG1 evaluates models for their ability to produce figures that clearly express and motivate the central idea of a paper (title, abstract, introduction, and figure caption) as input. Solving GENFIG1 requires more than producing visually appealing graphics: the task entails reasoning for text-to-image generation that couples scientific understanding with visual synthesis. Specifically, models must (i) comprehend and grasp the technical concepts of the paper, (ii) identify the most salient ones, and (iii) design a coherent and aesthetically effective graphic that conveys those concepts visually and is faithful to the input. We curate the benchmark from papers published at top deep-learning conferences, apply stringent quality control, and introduce an automatic evaluation metric that correlates well with expert human judgments. We evaluate a suite of representative models on GENFIG1 and demonstrate that the task presents significant challenges, even for the best-performing systems. We hope this benchmark serves as a foundation for future progress in multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。