arXiv:2505.11774cs.LGcs.AI2025-05NeurIPS被引 6

学生共建211道应用数学难题库,检验大模型真实推理能力。

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

  • 学生与教师协作设计难题,覆盖边界层分析等核心内容。
  • 前沿大模型在多数题目上表现不佳,暴露推理短板。
  • 互动迭代提升学生理解力,适合评估真实数学能力。

大型语言模型(LLMs)在数学求解方面进展显著,但评估多集中于具有精确解析解或形式证明的问题,忽略了应用科学与工程中普遍存在的近似类问题。为填补这一空白,我们基于前期工作,构建了HARDMath2数据集,包含211道原创题目,覆盖哈佛大学研究生应用数学课程的核心内容,包括边界层分析、WKB方法、非线性偏微分方程的渐近解,以及振荡积分的渐近行为。该数据集由课程师生共同设计并验证。通过一种创新协作环境,学生需根据课程大纲编写并优化难题,同行验证解法,测试不同模型,并自动比对大模型生成解与自身答案及数值真值。评估结果表明,当前领先模型在许多题目上仍表现不佳,揭示出大模型在数学推理上的差距。值得注意的是,学生通过与模型互动,识别其常见失败模式,逐步设计出更难的问题。这一双向过程不仅生成了更丰富、更具挑战性的基准,还显著提升了学生对课程内容的理解,这在顶尖语言模型可解决跨领域复杂问题的时代尤为重要。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To fill this gap, we build on prior work and present HARDMath2, a dataset of 211 original problems covering the core topics in an introductory graduate applied math class, including boundary-layer analysis, WKB methods, asymptotic solutions of nonlinear partial differential equations, and the asymptotics of oscillatory integrals. This dataset was designed and verified by the students and instructors of a core graduate applied mathematics course at Harvard. We build the dataset through a novel collaborative environment that challenges students to write and refine difficult problems consistent with the class syllabus, peer-validate solutions, test different models, and automatically check LLM-generated solutions against their own answers and numerical ground truths. Evaluation results show that leading frontier models still struggle with many of the problems in the dataset, highlighting a gap in the mathematical reasoning skills of current LLMs. Importantly, students identified strategies to create increasingly difficult problems by interacting with the models and exploiting common failure modes. This back-and-forth with the models not only resulted in a richer and more challenging benchmark but also led to qualitative improvements in the students' understanding of the course material, which is increasingly important as we enter an age where state-of-the-art language models can solve many challenging problems across a wide domain of fields.

应用数学大模型评测学生共建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。