arXiv:2505.00057cs.CL2025-05被引 1

评测大模型解高考数学题表现,揭示其教育应用潜力与短板。

A Report on the llms evaluating the high school questions

  • 用2019-2023年高考数学题测试8个大模型API
  • 模型答题准确率高但逻辑推理和创解能力仍有不足
  • 适合教育研究者、AI教学工具开发者参考

本报告旨在评估大语言模型(LLMs)在解答高中科学问题上的表现,并探索其在教育领域的潜在应用。随着自然语言处理中大语言模型的快速发展,其在教育中的应用已受到广泛关注。本研究选取2019至2023年高考数学试题作为评估数据,利用至少八个大语言模型API生成答案,并基于准确率、响应时间、逻辑推理与创造力等指标进行综合评估。通过深入分析评估结果,报告揭示了大模型在应对高中科学问题时的优势与局限,讨论了其对教育实践的影响。研究发现,尽管大模型在某些方面表现优异,但在逻辑推理与创造性问题解决能力上仍存在提升空间。该报告为大模型在教育领域的进一步研究与应用提供了实证基础,并提出了改进建议。

原文摘要 · Abstract (English)

This report aims to evaluate the performance of large language models (LLMs) in solving high school science questions and to explore their potential applications in the educational field. With the rapid development of LLMs in the field of natural language processing, their application in education has attracted widespread attention. This study selected mathematics exam questions from the college entrance examinations (2019-2023) as evaluation data and utilized at least eight LLM APIs to provide answers. A comprehensive assessment was conducted based on metrics such as accuracy, response time, logical reasoning, and creativity. Through an in-depth analysis of the evaluation results, this report reveals the strengths and weaknesses of LLMs in handling high school science questions and discusses their implications for educational practice. The findings indicate that although LLMs perform excellently in certain aspects, there is still room for improvement in logical reasoning and creative problem-solving. This report provides an empirical foundation for further research and application of LLMs in the educational field and offers suggestions for improvement.

大模型高考题教育应用推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。