arXiv:2410.14979cs.AIcs.CL2024-10中稿 · the CogSci 2025 co…被引 3

测试发现大模型解数学题靠直觉而非真推理

Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology

  • 用人类认知测试题改造数学题,检验大模型真实推理能力
  • 即使使用思维链提示,准确率仍下降50%以上
  • 错误答案显示模型依赖数据模式匹配,像人类直觉而非逻辑推理

大型语言模型(LLMs)解决数学问题的认知机制仍是未解之谜。目前缺乏可解释的实验证据将模型求解过程与人类认知心理学联系起来。为检验大模型是否具备类人数学推理能力,我们对人类认知反射测试(CRT)中的题目进行了改造。结果表明,即使采用思维链(CoT)提示,主流大模型(包括最新o1模型)在这些改造后的题目上错误率依然很高,平均准确率较原题下降高达50%。进一步分析其错误答案发现,模型主要依赖训练数据中的模式匹配,更接近人类直觉(系统1思维),而非人类逻辑推理(系统2思维)。这一发现挑战了大模型具备类人数学推理能力的普遍看法,可能促使人们重新评估其向通用人工智能演进的真实进展。

原文摘要 · Abstract (English)

The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable experimental evidence that connects LLMs' problem-solving with human cognitive psychology.To determine if LLMs possess human-like mathematical reasoning, we modified the problems used in the human Cognitive Reflection Test (CRT). Our results show that, even with the use of Chains of Thought (CoT) prompts, mainstream LLMs, including the latest o1 model (noted for its reasoning capabilities), have a high error rate when solving these modified CRT problems. Specifically, the average accuracy rate dropped by up to 50% compared to the original questions.Further analysis of LLMs' incorrect answers suggests that they primarily rely on pattern matching from their training data, which aligns more with human intuition (System 1 thinking) rather than with human-like reasoning (System 2 thinking). This finding challenges the belief that LLMs have genuine mathematical reasoning abilities comparable to humans. As a result, this work may adjust overly optimistic views on LLMs' progress towards artificial general intelligence.

认知科学大模型推理数学能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。