用定制GPT4辅助设计类作业评分,提升打分一致性和反馈质量。
The application of GPT-4 in grading design university students' assignment and providing feedback: An exploratory study
- 通过多轮迭代优化提示词,构建专用GPT模型以实现稳定评分。
- GPT评分一致性系数在0.65至0.78之间,达到教育评估可靠性标准。
- 适合教育工作者探索智能评分工具的开发与实践应用。
本研究探究GPT-4是否能有效评阅设计类大学生作业并提供有用反馈。设计类作业无唯一正确答案,常涉及开放式问题,不同背景教师评分存在差异。研究采用迭代方法开发定制GPT,旨在提升评分可靠性,并检验其能否提供建设性反馈。结果表明:经过多轮迭代,定制GPT与人工评分者间的一致性达到教育界普遍接受水平;GPT自身在不同时段评分的内部一致性(intra-reliability)介于0.65至0.78之间,符合教育评估对一致性和可比性的要求。研究最后验证了定制GPT可为学生提供有效反馈,并探讨了教师如何开发和迭代此类模型作为辅助评分工具。
原文摘要 · Abstract (English)
This study aims to investigate whether GPT-4 can effectively grade assignments for design university students and provide useful feedback. In design education, assignments do not have a single correct answer and often involve solving an open-ended design problem. This subjective nature of design projects often leads to grading problems,as grades can vary between different raters,for instance instructor from engineering background or architecture background. This study employs an iterative research approach in developing a Custom GPT with the aim of achieving more reliable results and testing whether it can provide design students with constructive feedback. The findings include: First,through several rounds of iterations the inter-reliability between GPT and human raters reached a level that is generally accepted by educators. This indicates that by providing accurate prompts to GPT,and continuously iterating to build a Custom GPT, it can be used to effectively grade students' design assignments, serving as a reliable complement to human raters. Second, the intra-reliability of GPT's scoring at different times is between 0.65 and 0.78. This indicates that, with adequate instructions, a Custom GPT gives consistent results which is a precondition for grading students. As consistency and comparability are the two main rules to ensure the reliability of educational assessment, this study has looked at whether a Custom GPT can be developed that adheres to these two rules. We finish the paper by testing whether Custom GPT can provide students with useful feedback and reflecting on how educators can develop and iterate a Custom GPT to serve as a complementary rater.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。