arXiv:2502.11574cs.AI2025-02被引 28

分析大模型数学推理缺陷,发现其常靠错误逻辑得正确答案

Large Language Models and Mathematical Reasoning Failures

  • 通过分析解题步骤而非仅看答案,定位模型推理漏洞
  • 8个主流模型在空间与策略推理上普遍出错,部分正确答案源于错误逻辑
  • 适合关注AI推理可信度的研究者与开发者

本文使用50道新构建的高中水平应用题,研究大语言模型(LLMs)的数学推理能力。不同于以往仅关注答案正确性的研究,我们严格分析最终答案与求解步骤,识别推理失败模式。评估了包括Mixtral、Llama、Gemini、GPT-4o及OpenAI o1系列在内的8个前沿模型,发现尽管较新模型(如o3-mini、deepseek-r1)准确率更高,但所有模型在空间推理、策略规划和算术运算上均存在错误,有时甚至以错误逻辑得出正确答案。常见失败模式包括无根据假设、过度依赖数字模式、难以将物理直觉转化为数学步骤。人工分析显示,模型在需要多步推导或现实知识的问题上表现不佳,尽管具备广泛数学知识。研究强调应评估推理过程而非仅答案,警示勿高估大模型的解题能力,并指出其泛化能力仍存显著缺口,亟需改进结构化推理与约束处理能力。

原文摘要 · Abstract (English)

This paper investigates the mathematical reasoning capabilities of large language models (LLMs) using 50 newly constructed high-school-level word problems. Unlike prior studies that focus solely on answer correctness, we rigorously analyze both final answers and solution steps to identify reasoning failures. Evaluating eight state-of-the-art models - including Mixtral, Llama, Gemini, GPT-4o, and OpenAI's o1 variants - we find that while newer models (e.g., o3-mini, deepseek-r1) achieve higher accuracy, all models exhibit errors in spatial reasoning, strategic planning, and arithmetic, sometimes producing correct answers through flawed logic. Common failure modes include unwarranted assumptions, over-reliance on numerical patterns, and difficulty translating physical intuition into mathematical steps. Manual analysis reveals that models struggle with problems requiring multi-step deduction or real-world knowledge, despite possessing broad mathematical knowledge. Our results underscore the importance of evaluating reasoning processes, not just answers, and caution against overestimating LLMs' problem-solving proficiency. The study highlights persistent gaps in LLMs' generalization abilities, emphasizing the need for targeted improvements in structured reasoning and constraint handling.

大模型数学推理推理错误结构化思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。