arXiv:2508.09932cs.AI2025-08被引 3

评测4大模型解数学题能力,发现推理错误常因步骤失误而非概念不清。

Mathematical Computation and Reasoning Errors by Large Language Models

  • 用自建难题测试模型解题过程,定位每步错误
  • o1模型准确率最高,双代理配置显著提升表现
  • 适合教育AI开发者和评估系统设计者参考

大型语言模型(LLMs)在数学教育中的教学与评估应用日益广泛。本研究评估了四种模型(OpenAI GPT-4o 和 o1、DeepSeek-V3 和 DeepSeek-R1)在算术、代数和数论三类数学任务上的表现,识别其解题过程中的逐步推理错误。研究不依赖标准基准,而是通过项目模型构建对LLM具有挑战性的题目。系统分析了最终答案准确率及各解题步骤的错误情况,测试了单代理与双代理配置。结果显示,推理增强型 OpenAI o1 模型在三类任务中均达到高或近乎完美的准确率;错误分析表明,程序性失误最常见且显著影响整体表现,而概念性误解较少。双代理配置能显著提升整体性能。研究为提升模型表现提供了可操作建议,并支持将LLM更可靠地融入数学教育场景。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly utilized in AI-driven educational instruction and assessment, particularly within mathematics education. The capability of LLMs to generate accurate answers and detailed solutions for math problem-solving tasks is foundational for ensuring reliable and precise feedback and assessment in math education practices. Our study focuses on evaluating the accuracy of four LLMs (OpenAI GPT-4o and o1, DeepSeek-V3 and DeepSeek-R1) solving three categories of math tasks, including arithmetic, algebra, and number theory, and identifies step-level reasoning errors within their solutions. Instead of relying on standard benchmarks, we intentionally build math tasks (via item models) that are challenging for LLMs and prone to errors. The accuracy of final answers and the presence of errors in individual solution steps were systematically analyzed and coded. Both single-agent and dual-agent configurations were tested. It is observed that the reasoning-enhanced OpenAI o1 model consistently achieved higher or nearly perfect accuracy across all three math task categories. Analysis of errors revealed that procedural slips were the most frequent and significantly impacted overall performance, while conceptual misunderstandings were less frequent. Deploying dual-agent configurations substantially improved overall performance. These findings offer actionable insights into enhancing LLM performance and underscore effective strategies for integrating LLMs into mathematics education, thereby advancing AI-driven instructional practices and assessment precision.

数学推理大模型评测教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。