首个专评大模型数学绘图能力的基准,覆盖文本到代码和图像生成。
Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

- 构建融合文本到代码与图像生成的统一评估框架。
- 从2920个竞赛题中筛选高质量含视觉信息的数学图示数据。
- 揭示当前大模型在数学绘图任务上表现普遍不佳。
从文本提示生成数学精确图示已成为大型语言模型(LLMs)一项关键但研究不足的能力,对课程设计、题目集自动评分和科学出版等领域具有重要意义。实现该能力需空间推理、数学推理与渲染系统的精准协同。现有基准如MathVision、MathVista专注于数学推理,DiagramGenBenchmark关注通用绘图,MermaidSeqBench则聚焦通用图表生成,但均未提供专门用于评估数学图示生成的标准化提示-图像对。本文提出Math-Vision Diagrams,首个专为评估LLMs数学图示生成能力设计的基准,首次在统一设置下同时评估文本到代码与文本到图像生成范式,且不依赖特定编程语言或模型类型。基于Math-Vision基准,我们从3040个样本中精选出2920个来自高质量竞赛题目的图像,这些题目包含必要的视觉上下文。提出一种结合多模型集成与领域专家(SME)人工校验的新颖数据构建流程,并配套评估指标。对多个领先模型的测试表明,当前大模型在数学图示生成任务上仍存在明显不足。所有代码、数据、标注流程及评估脚本将完全开源。
原文摘要 · Abstract (English)
The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs). This has been of interest to researchers in the areas of curriculum preparation, automated ranking of problem sets, and scientific publishing. For LLMs to achieve this, it requires per- fect coordination between Spatial Reasoning, Mathematical Reasoning, and Rendering systems. While existing benchmarks such as MathVision, MathVista are built for Math Reasoning or DiagramGenBenchmark, Mer- maidSeqBench on general purpose diagram generation, no prior work provides a standardized set of prompt, image pairs that can be used to evaluate the LLMs specifically on math diagram generation. This includes fields that span both both text-to-code and text-to-image paradigms. We introduce Math-Vision Diagrams, the first benchmark specifically designed to evaluate LLMs on mathematical diagram generation, and the first to assess text-to-code and text-to-image generation paradigms together in a single unified setting, agnostic of the underlying coding lan- guage or model type. Building on the Math-Vision benchmark, we select a subset of 2920 images out of 3040 from high-quality competition problems with essential visual context. A novel pipeline combining an ensemble of LLMs with Subject Matter Expert (SME) curation is presented, together with a suite of evaluation metrics. Testing several leading models against this benchmark, we demonstrate that LLMs struggle with math diagram generation. All code, data, curation pipeline, and evaluation scripts will be fully open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。