对比三款大模型解数学题能力,发现GPT-4o最稳定,尤其擅长高阶题目。
Evaluation of LLMs for mathematical problem solving
- 基于五维评估框架,系统分析推理过程与答案质量
- GPT-4o在复杂题集上表现最优,准确率领先
- 各模型短板清晰:解释不足、漏步骤、逻辑僵化
大型语言模型(LLMs)在教育任务中表现优异,但在数学问题求解方面仍研究不足。本研究对比了GPT-4o、DeepSeek-V3和Gemini-2.0三款主流模型,在GSM8K、MATH500及MIT开放课程数据集上的表现。采用基于结构化思维链(SCoT)的五维评估框架,从最终答案正确性、步骤完整性、步骤有效性、中间计算准确性及问题理解力五个维度进行分析。结果表明,GPT-4o在所有数据集上均表现最稳定,尤其在MIT开放课程数据集的高阶问题中优势显著;DeepSeek-V3在优化等结构化任务中表现良好,但在统计推断任务中波动较大;Gemini-2.0虽语言表达清晰,但多步推理与符号逻辑能力较弱。错误分析显示:GPT-4o解释不足,DeepSeek-V3遗漏中间步骤,Gemini-2.0在高维数学推理中灵活性差。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive performance on a range of educational tasks, but are still understudied for their potential to solve mathematical problems. In this study, we compare three prominent LLMs, including GPT-4o, DeepSeek-V3, and Gemini-2.0, on three mathematics datasets of varying complexities (GSM8K, MATH500, and MIT Open Courseware datasets). We take a five-dimensional approach based on the Structured Chain-of-Thought (SCoT) framework to assess final answer correctness, step completeness, step validity, intermediate calculation accuracy, and problem comprehension. The results show that GPT-4o is the most stable and consistent in performance across all the datasets, but particularly it performs outstandingly in high-level questions of the MIT Open Courseware dataset. DeepSeek-V3 is competitively strong in well-structured domains such as optimisation, but suffers from fluctuations in accuracy in statistical inference tasks. Gemini-2.0 shows strong linguistic understanding and clarity in well-structured problems but performs poorly in multi-step reasoning and symbolic logic. Our error analysis reveals particular deficits in each model: GPT-4o is at times lacking in sufficient explanation or precision; DeepSeek-V3 leaves out intermediate steps; and Gemini-2.0 is less flexible in mathematical reasoning in higher dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。