arXiv:2503.14630cs.SEcs.AI2025-03被引 7

评测大模型生成编程反馈的准确性与可靠性

Assessing Large Language Models for Automated Feedback Generation in Learning Programming Problem Solving

  • 对比4个大模型在45份学生代码上生成反馈
  • 63%反馈准确完整,37%存在错误或幻觉
  • 适合教育科技开发者和编程教学研究者参考

在编程问题求解学习中,提供有效反馈对学生成长至关重要。在此背景下,大型语言模型(LLMs)作为自动化反馈生成的潜在工具应运而生。然而,它们在识别学生代码中推理错误方面的可靠性与能力仍不明确。本研究在包含45份学生解答的基准数据集上评估了四种LLM(GPT-4o、GPT-4o mini、GPT-4-Turbo 和 Gemini-1.5-pro)的表现,重点考察其提供准确且有洞见反馈的能力,尤其是对推理错误的识别。分析显示,63%的反馈提示准确且完整,而37%存在错误,包括错误行号定位、解释不当或虚构问题。这些发现揭示了大模型在编程教育中的潜力与局限,强调需进一步提升其可靠性,以降低教育应用中的风险。

原文摘要 · Abstract (English)

Providing effective feedback is important for student learning in programming problem-solving. In this sense, Large Language Models (LLMs) have emerged as potential tools to automate feedback generation. However, their reliability and ability to identify reasoning errors in student code remain not well understood. This study evaluates the performance of four LLMs (GPT-4o, GPT-4o mini, GPT-4-Turbo, and Gemini-1.5-pro) on a benchmark dataset of 45 student solutions. We assessed the models' capacity to provide accurate and insightful feedback, particularly in identifying reasoning mistakes. Our analysis reveals that 63\% of feedback hints were accurate and complete, while 37\% contained mistakes, including incorrect line identification, flawed explanations, or hallucinated issues. These findings highlight the potential and limitations of LLMs in programming education and underscore the need for improvements to enhance reliability and minimize risks in educational applications.

大模型编程教育反馈生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。