用冷门数学竞赛题测试大模型,发现几何题是短板。
Evaluating the Reasoning Abilities of LLMs on Underrepresented Mathematics Competition Problems
- 用密苏里大学数学竞赛题评估三款大模型的解题能力。
- 深求-V3表现最好,但几何类题目整体准确率低。
- 不同模型错因各异,几何推理仍是主要难点。
近期多项研究关注大语言模型(LLMs)在数学推理方面的局限性,但多数使用相同数据集,限制了结论的普适性。本研究聚焦于未被充分覆盖的数学竞赛题,对GPT-4o-mini、Gemini-2.0-Flash和DeepSeek-V3三款主流模型进行评估,采用密苏里大学数学竞赛中的微积分、解析几何与离散数学题目。通过对比模型回答与标准答案,分析其准确率与推理过程。结果显示,DeepSeek-V3在三个领域均表现最佳;所有模型在几何类问题上表现明显薄弱。其中,DeepSeek-V3主要出错于计算与逻辑失误,GPT-4o-mini多因逻辑错误或方法不当,Gemini则常出现推理不完整及过早下结论。研究表明,在非主流数学竞赛数据集上评估有助于揭示模型差异化的错误模式,凸显结构化推理尤其是几何领域的持续挑战。
原文摘要 · Abstract (English)
Understanding the limitations of Large Language Models, or LLMs, in mathematical reasoning has been the focus of several recent studies. However, the majority of these studies use the same datasets for benchmarking, which limits the generalizability of their findings and may not fully capture the diverse challenges present in mathematical tasks. The purpose of the present study is to analyze the performance of LLMs on underrepresented mathematics competition problems. We prompted three leading LLMs, namely GPT-4o-mini, Gemini-2.0-Flash, and DeepSeek-V3, with the Missouri Collegiate Mathematics Competition problems in the areas of Calculus, Analytic Geometry, and Discrete Mathematics. The LLMs responses were then compared to the known correct solutions in order to determine the accuracy of the LLM for each problem domain. We also analyzed the LLMs reasoning to explore patterns in errors across problem types and models. DeepSeek-V3 has the best performance in all three categories of Calculus, Analytic Geometry, and Discrete Mathematics, both in reasoning and correct final answers. All three LLMs exhibited notably weak performance in Geometry. The majority of errors made by DeepSeek-V3 were attributed to computational and logical mistakes, whereas GPT-4o-mini frequently exhibited logical and approach-related errors. Gemini, on the other hand, tended to struggle with incomplete reasoning and drawing rushed conclusions. In conclusion, evaluating LLMs on underrepresented mathematics competition datasets can provide deeper insights into their distinct error patterns and highlight ongoing challenges in structured reasoning, particularly within the domain of Geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。