AI辅助批改论文,能高效发现评分差异,但需人工复核。
Assessing instructor-AI cooperation for grading essay-type questions in an introductory sociology course
- 用GPT模型自动转写和评分70份手写答卷,测试不同设置下的表现。
- GPT评分与人工评分高度相关,模板答案可提升准确性。
- 适合教师用于快速筛查不一致评分,提升批改公平性与效率。
本研究探讨人工智能(AI)作为高等教育中作文题评分的辅助工具,重点考察其与人工评分的一致性及降低偏见的潜力。基于一门入门社会学课程的70份手写试卷,评估了生成式预训练变换器(GPT)模型在文本转录和评分任务中的表现。GPT模型在不同设置下进行了转录与评分测试。结果显示,人工与GPT转录高度一致,GPT-4o-mini在准确率上优于GPT-4o;评分方面,当提供标准答案模板时,GPT评分与人工评分呈现强相关性。然而,仍存在差异,表明GPT更适合作为“第二评审者”来标记评分不一致,而非完全替代人工评估。本研究为教育中AI应用提供了实证支持,展示了其在提升作文评分公平性与效率方面的潜力。
原文摘要 · Abstract (English)
This study explores the use of artificial intelligence (AI) as a complementary tool for grading essay-type questions in higher education, focusing on its consistency with human grading and potential to reduce biases. Using 70 handwritten exams from an introductory sociology course, we evaluated generative pre-trained transformers (GPT) models' performance in transcribing and scoring students' responses. GPT models were tested under various settings for both transcription and grading tasks. Results show high similarity between human and GPT transcriptions, with GPT-4o-mini outperforming GPT-4o in accuracy. For grading, GPT demonstrated strong correlations with the human grader scores, especially when template answers were provided. However, discrepancies remained, highlighting GPT's role as a "second grader" to flag inconsistencies for assessment reviewing rather than fully replace human evaluation. This study contributes to the growing literature on AI in education, demonstrating its potential to enhance fairness and efficiency in grading essay-type questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。