测试大模型在数学推理中的基本能力,发现其仅能解决确定性问题。
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
- 通过逐步升级难度的数学题测试模型表现
- 90%以上准确率仅限于基础运算与判定类任务
- 面对组合搜索难题时完全失效,暴露其缺乏创造性思维
大型语言模型虽展现强大涌现能力,但其数值推理的鲁棒性仍存疑。本文通过100道递进式数学题评估多个先进模型,涵盖四类任务:(1) 基础算术,(2) 高级运算,(3) 质数判断,(4) 24点数谜游戏。结果显示,模型在前三类任务中准确率超过90%,表现出对确定性算法的良好执行能力;但在需要大规模组合搜索的24点数谜中普遍失败,揭示其无法进行生成式求解。这表明模型的所谓数学推理更多依赖模式匹配而非灵活分析,限制其在需创新思维的数值任务中的应用潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable emergent capabilities, yet the robustness of their numerical reasoning remains an open question. While standard benchmarks evaluate LLM reasoning on complex problem sets using aggregated metrics, they often obscure foundational weaknesses. In this work, we probe LLM mathematical numeracy by evaluating performance on problems of escalating complexity, from constituent operations to combinatorial puzzles. We test several state-of-the-art LLM-based agents on a 100-problem challenge comprising four categories: (1) basic arithmetic, (2) advanced operations, (3) primality checking, and (4) the Game of 24 number puzzle. Our results show that while the agents achieved high accuracy on the first three categories, which require deterministic algorithmic execution, they consistently failed at the number puzzle, underlining its demand for a heuristic search over a large combinatorial space to be a significant bottleneck. These findings reveal that the agents' proficiency is largely confined to recalling and executing known algorithms, rather than performing generative problem-solving. This suggests their apparent numerical reasoning is more akin to sophisticated pattern-matching than flexible, analytical thought, limiting their potential for tasks that require novel or creative numerical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。