新数据集评测大模型高阶数学能力,发现其仍难应对研究生级计算数学问题。
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
- 自创1500道研究生级数学多选题,覆盖线代、优化、概率等核心领域。
- 顶尖大模型在该数据集上准确率不足50%,表明高阶数学推理仍是难点。
- 专为客观评估设计,适合研究大模型数学能力与教育应用的学者使用。
大型语言模型(LLMs)在数学推理方面的评估主要集中在初等数学、竞赛题或形式化定理证明,而对研究生级和计算数学的关注相对不足。本文提出CompMath-MCQ,一个用于评估大模型在多选题形式下高级数学推理能力的新基准数据集。该数据集包含1,500道由研究生课程教授原创的题目,涵盖线性代数、数值优化、向量微积分、概率及基于Python的科学计算等主题。每道题提供三个选项,仅一个正确。为防止数据泄露,所有题目均为新创,未来自已有材料。通过跨模型意见分歧检测结合人工专家审核确保题目有效性。采用多选格式,支持通过lm_eval库进行客观、可复现、无偏见的评估。现有最先进大模型的基线结果表明,高级计算数学推理仍是重大挑战。数据集已开源:https://github.com/biancaraimondi/CompMath-MCQ.git。
原文摘要 · Abstract (English)
The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving graduate-level and computational mathematics relatively underexplored. We introduce CompMath-MCQ, a new benchmark dataset for assessing LLMs on advanced mathematical reasoning in a multiple-choice setting. The dataset consists of 1{,}500 originally authored questions by professors of graduate-level courses, covering topics including Linear Algebra, Numerical Optimization, Vector Calculus, Probability, and Python-based scientific computing. Three option choices are provided for each question, with exactly one of them being correct. To ensure the absence of data leakage, all questions are newly created and not sourced from existing materials. The validity of questions is verified through a procedure based on cross-LLM disagreement, followed by manual expert review. By adopting a multiple-choice format, our dataset enables objective, reproducible, and bias-free evaluation through lm_eval library. Baseline results with state-of-the-art LLMs indicate that advanced computational mathematical reasoning remains a significant challenge. We release CompMath-MCQ at the following link: https://github.com/biancaraimondi/CompMath-MCQ.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。