构建研究生级应用数学难题数据集,揭示大模型在此类问题上的严重短板。
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
- 基于渐近分析课程自动生成带验证解的高难度应用数学题
- 顶尖模型GPT-4在366题测试集上准确率仅43.8%
- 适合研究大模型数学推理能力或评估其在科学建模中的适用性
现有大语言模型(LLM)基准数据集对高级应用数学问题覆盖不足。为此,我们引入HARDMath,该数据集源自渐近方法研究生课程,包含需解析近似技巧的高难度应用数学题。这些问题结合了数学推理、计算工具与主观判断,对LLM构成挑战。我们的框架可自动生成大量问题,并通过数值真值验证解的正确性。我们在HARDMath-mini(366道题的子采样测试集)及40道应用科学情境下的文字题上评估了开源与闭源模型。即使使用少样本思维链提示,顶级闭源模型GPT-4的整体准确率也仅为43.8%,且所有模型表现均显著低于现有数学基准数据集。我们还进行了详细的错误分析,揭示了模型失败的关键模式。结果表明当前LLM在高等研究生级应用数学问题上存在明显局限,凸显HARDMath类数据集对提升模型数学能力的重要性。
原文摘要 · Abstract (English)
Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce HARDMath, a dataset inspired by a graduate course on asymptotic methods, featuring challenging applied mathematics problems that require analytical approximation techniques. These problems demand a combination of mathematical reasoning, computational tools, and subjective judgment, making them difficult for LLMs. Our framework auto-generates a large number of problems with solutions validated against numerical ground truths. We evaluate both open- and closed-source LLMs on HARDMath-mini, a sub-sampled test set of 366 problems, as well as on 40 word problems formulated in applied science contexts. Even leading closed-source models like GPT-4 achieve only 43.8% overall accuracy with few-shot Chain-of-Thought prompting, and all models demonstrate significantly lower performance compared to results on existing mathematics benchmark datasets. We additionally conduct a detailed error analysis to gain insights into the failure cases of LLMs. These results demonstrate limitations of current LLM performance on advanced graduate-level applied math problems and underscore the importance of datasets like HARDMath to advance mathematical abilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。