对比5大模型解微积分题,发现最强模型正确率超94%。
Benchmarking Large Language Models for Calculus Problem-Solving: A Comparative Analysis
- 用互相出题的方式测试模型解题能力
- ChatGPT 4o正确率最高达94.71%,Meta AI仅56.75%
- 模型在概念理解与代数运算上仍有明显短板
本研究对五款主流大语言模型(ChatGPT 4o、Copilot Pro、Gemini Advanced、Claude Pro 和 Meta AI)在微积分求导问题上的表现进行了全面评估。通过13种基础题型的系统性交叉评测,每个模型需解答其他模型生成的问题。结果显示性能差异显著:ChatGPT 4o 成功率达94.71%,居首;Claude Pro 为85.74%;Gemini Advanced 为84.42%;Copilot Pro 为76.30%;Meta AI 仅为56.75%。所有模型在程序化求导任务中表现良好,但在概念理解与代数处理方面存在局限。尤其是涉及单调区间与优化应用题时,各模型均面临挑战。交叉评测矩阵显示,Claude Pro 生成的问题最难,表明其生成能力与求解能力存在差异。研究结果对教育应用具有重要意义:尽管模型具备强大程序处理能力,但其概念理解仍远不及人类数学推理,凸显了人工教学在培养深层数学思维中的不可替代性。
原文摘要 · Abstract (English)
This study presents a comprehensive evaluation of five leading large language models (LLMs) - Chat GPT 4o, Copilot Pro, Gemini Advanced, Claude Pro, and Meta AI - on their performance in solving calculus differentiation problems. The investigation assessed these models across 13 fundamental problem types, employing a systematic cross-evaluation framework where each model solved problems generated by all models. Results revealed significant performance disparities, with Chat GPT 4o achieving the highest success rate (94.71%), followed by Claude Pro (85.74%), Gemini Advanced (84.42%), Copilot Pro (76.30%), and Meta AI (56.75%). All models excelled at procedural differentiation tasks but showed varying limitations with conceptual understanding and algebraic manipulation. Notably, problems involving increasing/decreasing intervals and optimization word problems proved most challenging across all models. The cross-evaluation matrix revealed that Claude Pro generated the most difficult problems, suggesting distinct capabilities between problem generation and problem-solving. These findings have significant implications for educational applications, highlighting both the potential and limitations of LLMs as calculus learning tools. While they demonstrate impressive procedural capabilities, their conceptual understanding remains limited compared to human mathematical reasoning, emphasizing the continued importance of human instruction for developing deeper mathematical comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。