arXiv:2506.06390cs.CYcs.AI2025-06被引 9

评测大模型在电路分析作业批改中的表现,发现GPT-4o和Llama 3更优。

Benchmarking Large Language Models on Homework Assessment in Circuit Analysis

  • 构建电路分析题的LaTeX格式学生作答数据集,支持五维评估
  • GPT-4o与Llama 3 70B在完整性、答案正确性等指标上显著优于GPT-3.5 Turbo
  • 研究结果可为智能教学系统开发提供可靠基准,适合工程教育研究者

大语言模型(LLMs)在代码开发、机器人、金融和教育等领域具有变革潜力,因其具备广泛先验知识并持续进步。本文研究了LLMs在工程教育中的应用,具体评估GPT-3.5 Turbo、GPT-4o和Llama 3 70B在本科电路分析课程作业批改中的能力。我们构建了一个包含官方参考解答和真实学生解答的新数据集,覆盖多个电路分析主题。为克服当前主流LLMs在图像识别上的局限,所有解答均转换为LaTeX格式。基于该数据集,设计提示模板以测试学生解答的五个维度:完整性、解题方法、最终答案、算术错误和单位。结果显示,GPT-4o和Llama 3 70B在全部五项指标上显著优于GPT-3.5 Turbo,且两者在不同评估方面各有优势。此外,本文还揭示了当前LLMs在电路分析中若干方面的局限性。鉴于确保LLM生成批改结果可靠性对避免误导学生至关重要,本研究建立了基准,并为未来开发可靠的个性化电路分析辅导系统提供了重要洞见。所提出的评估方法亦可推广至其他工程类课程。

原文摘要 · Abstract (English)

Large language models (LLMs) have the potential to revolutionize various fields, including code development, robotics, finance, and education, due to their extensive prior knowledge and rapid advancements. This paper investigates how LLMs can be leveraged in engineering education. Specifically, we benchmark the capabilities of different LLMs, including GPT-3.5 Turbo, GPT-4o, and Llama 3 70B, in assessing homework for an undergraduate-level circuit analysis course. We have developed a novel dataset consisting of official reference solutions and real student solutions to problems from various topics in circuit analysis. To overcome the limitations of image recognition in current state-of-the-art LLMs, the solutions in the dataset are converted to LaTeX format. Using this dataset, a prompt template is designed to test five metrics of student solutions: completeness, method, final answer, arithmetic error, and units. The results show that GPT-4o and Llama 3 70B perform significantly better than GPT-3.5 Turbo across all five metrics, with GPT-4o and Llama 3 70B each having distinct advantages in different evaluation aspects. Additionally, we present insights into the limitations of current LLMs in several aspects of circuit analysis. Given the paramount importance of ensuring reliability in LLM-generated homework assessment to avoid misleading students, our results establish benchmarks and offer valuable insights for the development of a reliable, personalized tutor for circuit analysis -- a focus of our future work. Furthermore, the proposed evaluation methods can be generalized to a broader range of courses for engineering education in the future.

大模型评估教育应用电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。