用精心设计的提示让ChatGPT高效批改编程作业,降低错误率。
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant
- 通过上下文学习设计提示,自动分析反馈结构并评估性能。
- GPT-4T纠错能力显著优于GPT-3.5T,但仍有误报风险。
- 适合教育技术研究者和编程教学工具开发者参考。
大型语言模型(LLMs)使师生比达到1:1成为可能。本文聚焦于利用LLMs作为大学入门编程课程的教学助教,评估GPT-3.5T与GPT-4T在提供学生反馈方面的表现。实证结果表明,GPT-4T远优于GPT-3.5T,但仍不适合直接投入实际使用,因存在生成错误信息的风险且用户难以察觉。为此,本文提出一种基于上下文学习的精细提示设计,可自动化评估过程,并给出错误反馈比例的下界,节省人力。该方法生成的反馈具备可程序分析的结构,包含模型解决任务时的诊断信息。此外,文章还提出了一种基于所提提示技术的实用教学工具实现策略,从教学法角度开辟了新的可能性。
原文摘要 · Abstract (English)
The dream of achieving a student-teacher ratio of 1:1 is closer than ever thanks to the emergence of large language models (LLMs). One potential application of these models in the educational field would be to provide feedback to students in university introductory programming courses, so that a student struggling to solve a basic implementation problem could seek help from an LLM available 24/7. This article focuses on studying three aspects related to such an application. First, the performance of two well-known models, GPT-3.5T and GPT-4T, in providing feedback to students is evaluated. The empirical results showed that GPT-4T performs much better than GPT-3.5T, however, it is not yet ready for use in a real-world scenario. This is due to the possibility of generating incorrect information that potential users may not always be able to detect. Second, the article proposes a carefully designed prompt using in-context learning techniques that allows automating important parts of the evaluation process, as well as providing a lower bound for the fraction of feedbacks containing incorrect information, saving time and effort. This was possible because the resulting feedback has a programmatically analyzable structure that incorporates diagnostic information about the LLM's performance in solving the requested task. Third, the article also suggests a possible strategy for implementing a practical learning tool based on LLMs, which is rooted on the proposed prompting techniques. This strategy opens up a whole range of interesting possibilities from a pedagogical perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。