测试大模型大学数学能力,含图文混合题
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
- 构建1100道未公开的大学级数学题库,覆盖六大学科
- 图文题模型准确率仅58.5%,远低于纯文本题的93.1%
- 引入μ-MATH数据集评估模型判题能力,顶尖模型F1仅90.1%
当前大模型数学能力评估存在局限:现有基准规模小,多聚焦中小学题目,且主题单一。此外,任务中视觉元素的应用仍不充分。为弥补这些不足,我们提出U-MATH,一个包含1,100道未公开的大学级开放性问题的新基准,源自教学材料,涵盖六大学科,其中20%为多模态问题。由于题型为开放作答,我们采用大模型来评判生成解法的正确性,并发布μ-MATH数据集以评估模型的判题能力。对领先大模型的评测显示,其在多模态推理上存在显著短板:文本任务最高准确率达93.1%,而视觉任务仅为58.5%。此外,解法评判极具挑战,最先进模型虽达最高性能,但F1分数仍仅90.1%。
原文摘要 · Abstract (English)
The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lack diversity in topics. Additionally, the inclusion of visual elements in tasks remains largely under-explored. To address these gaps, we introduce U-MATH, a novel benchmark of 1,100 unpublished open-ended university-level problems sourced from teaching materials. It is balanced across six core subjects, with 20% of multimodal problems. Given the open-ended nature of U-MATH problems, we employ an LLM to judge the correctness of generated solutions. To this end, we release $μ$-MATH, a dataset to evaluate the LLMs' capabilities in judging solutions. Benchmarking leading LLMs reveals marked limitations in multi-modal reasoning, with maximum accuracy reaching 93.1\% on textual tasks but only 58.5\% on visual ones. Furthermore, solution judgment proves challenging, requiring the most advanced models to achieve meaningfully high performance, even still peaking at an imperfect F1-score of 90.1\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。