arXiv:2510.08595cs.CL2025-10

发现大模型数学推理存在严重脆弱性,能算但不会变通。

Systematic Diagnosis of Brittle Reasoning in Large Language Models

  • 用GPT-3.5生成分步推理,GPT-4o-mini分析错误并聚类出不同推理模式。
  • 在顺序计算上准确率接近完美,但组合推理受限时性能骤降。
  • 揭示模型非人类的脆弱认知结构,适合关注AI可靠性研究者。

人工智能的核心问题之一是机器学习模型对数学的理解程度。为此,我们提出一种新型数学推理评估框架,超越传统基准,诊断具体失败点。方法首先利用gpt-3.5-turbo在GSM8K数据集上生成结构化分步推理,再通过更强大的analyst模型gpt-4o-mini对错误进行分类,并关键性地对每个推理句进行无监督聚类,识别出涌现的“推理模式”。分析揭示了一种非人般的脆弱认知特征:尽管在顺序计算等程序化模式上达到近完美准确率,但在需要受限制的组合推理模式下表现急剧下滑。通过识别并量化这些不同推理能力的可靠性,本工作提供了更细粒度的数学理解评估方法,并为未来能力提升与应用可靠性改进提供了精确路线图。

原文摘要 · Abstract (English)

A central question in artificial intelligence is the extent to which machine learning models comprehend mathematics. To address this, we propose a novel framework for measuring mathematical reasoning that moves beyond standard benchmarks to diagnose specific failure points. Our method first generates structured, step-by-step reasoning from gpt-3.5-turbo on the GSM8K dataset. We then use a more capable analyst model, gpt-4o-mini, to categorize errors and, crucially, perform an unsupervised clustering of every reasoning sentence to identify emergent "reasoning modes." This analysis reveals a cognitive profile with a stark, nonhuman-like brittleness: while the model achieves near-perfect accuracy on procedural modes like sequential calculation, its performance on modes requiring combinatorial reasoning with restrictions plummets. By identifying and quantifying the reliability of these distinct reasoning skills, our work provides a more granular method to evaluate mathematical comprehension and offers a precise roadmap for developing new capabilities and more reliable future applications.

大模型数学推理脆弱性认知分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。