对比7大模型解高数题,发现准确率差异大,重试提示能救错答案。
Performance Comparison of Large Language Models on Advanced Calculus Problems
- 用32道题(共320分)测试7个大模型的解题能力
- ChatGPT 4o和Mistral AI表现稳定,Gemini和Meta AI在积分优化题出错多
- 重试提示可纠正错误,对教学与应用有指导意义
本文深入分析了七种大型语言模型(LLMs)在解决多样化高等数学微积分问题上的表现。研究评估了ChatGPT 4o、Gemini Advanced with 1.5 Pro、Copilot Pro、Claude 3.5 Sonnet、Meta AI、Mistral AI和Perplexity的准确性、可靠性及解题能力。测试包含三十二道题目,总计320分,涵盖向量计算、几何解释、积分求解和优化任务。结果揭示显著趋势:ChatGPT 4o和Mistral AI在各类问题中保持一致高准确率,体现其数学求解的稳健性;而Gemini Advanced with 1.5 Pro和Meta AI在复杂积分与优化问题中表现较弱,表明改进空间。研究还强调重提示的重要性——多个案例显示模型初答错误,经重提示后修正。该研究为教育、科研及开发人员提供了大模型在数学领域能力与局限的全面理解,推动其技术迭代。
原文摘要 · Abstract (English)
This paper presents an in-depth analysis of the performance of seven different Large Language Models (LLMs) in solving a diverse set of math advanced calculus problems. The study aims to evaluate these models' accuracy, reliability, and problem-solving capabilities, including ChatGPT 4o, Gemini Advanced with 1.5 Pro, Copilot Pro, Claude 3.5 Sonnet, Meta AI, Mistral AI, and Perplexity. The assessment was conducted through a series of thirty-two test problems, encompassing a total of 320 points. The problems covered various topics, from vector calculations and geometric interpretations to integral evaluations and optimization tasks. The results highlight significant trends and patterns in the models' performance, revealing both their strengths and weaknesses - for instance, models like ChatGPT 4o and Mistral AI demonstrated consistent accuracy across various problem types, indicating their robustness and reliability in mathematical problem-solving, while models such as Gemini Advanced with 1.5 Pro and Meta AI exhibited specific weaknesses, particularly in complex problems involving integrals and optimization, suggesting areas for targeted improvements. The study also underscores the importance of re-prompting in achieving accurate solutions, as seen in several instances where models initially provided incorrect answers but corrected them upon re-prompting. Overall, this research provides valuable insights into the current capabilities and limitations of LLMs in the domain of math calculus, with the detailed analysis of each model's performance on specific problems offering a comprehensive understanding of their strengths and areas for improvement, contributing to the ongoing development and refinement of LLM technology. The findings are particularly relevant for educators, researchers, and developers seeking to leverage LLMs for educational and practical applications in mathematics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。