arXiv:2504.00226cs.AI2025-04被引 4

测试大模型对数字的基本理解能力,发现其在复杂计算中表现脆弱。

Large Language Models in Numberland: A Quick Test of Their Numerical Reasoning Abilities

  • 设计100题的'Numberland'测试,评估大模型基础数感与计算整合能力。
  • 基础运算正确率74%-95%,但24点游戏仅10%-73%,搜索成主要瓶颈。
  • 顶尖模型在更难题目上准确率暴跌至27%,揭示其数学推理的局限性。

人类数学推理的关键在于数感——对数字及其关系的抽象理解,使我们能用有限资源处理庞大的数字空间。尽管大语言模型(LLMs)常在高阶数学问题(如奥数、几何、应用题和谜题)中被评测,但其底层数感仍较少被探索。本文提出‘Numberland’测试集,包含100道题目,涵盖基本运算、高级计算(如幂运算、复数)、质数判断及24点游戏,旨在检验基础技能及其在复杂不确定问题中的整合能力。我们评估了五款基于LLM的智能体:OpenAI的o1和o1-mini、Google Gemini、Microsoft Copilot以及Anthropic Claude。在前三类允许确定性解法的任务中,它们得分74%-95%;但在需试错搜索的24点游戏中,性能降至10%-73%。我们进一步对表现最优的o1模型(24点准确率73%)在25个更难题目上测试,其得分降至27%,证实搜索是核心瓶颈。错误类型分析显示,大模型的数理推理存在脆弱性,令人意外。这些结果表明,简单的专项测试可有效揭示并解释大模型数学能力的边界,有助于保障其安全应用。

原文摘要 · Abstract (English)

An essential element of human mathematical reasoning is our number sense -- an abstract understanding of numbers and their relationships -- which allows us to solve problems involving vast number spaces using limited computational resources. Mathematical reasoning of Large Language Models (LLMs) is often tested on high-level problems (such as Olympiad challenges, geometry, word problems, and puzzles), but their low-level number sense remains less explored. We introduce "Numberland," a 100-problem test to evaluate the numerical reasoning abilities of LLM-based agents. The tasks -- basic operations, advanced calculations (e.g., exponentiation, complex numbers), prime number checks, and the 24 game -- aim to test elementary skills and their integration in solving complex and uncertain problems. We evaluated five LLM-based agents: OpenAI's o1 and o1-mini, Google Gemini, Microsoft Copilot, and Anthropic Claude. They scored 74-95% on the first three tasks that allow deterministic steps to solutions. In the 24 game, which needs trial-and-error search, performance dropped to 10-73%. We tested the top 24 solver (o1 with 73% accuracy) on 25 harder problems, and its score fell to 27%, confirming search as a bottleneck. These results, along with the types of mistakes, suggest a fragile number of LLMs, which is a bit surprising given their prowess in challenging benchmarks. The limits of LLM numerical reasoning highlight the scope of simple, targeted tests to evaluate and explain LLM math skills to ensure safe use.

大模型数学推理数感测试24点游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。