通过要求学生反思答题思路,让AI反馈真正提升学习效果。
How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors
- 让学生解释答案并打信心分,引导深度思考。
- 40%学生按反馈去读教材,远高于传统阅读率。
- 学生报告自信心和知识点应用能力明显提升。
在大型本科理工课程中,规模化提供个性化、详细的反馈仍是一大挑战。我们设计并实证评估了一套实践测验系统,将AI生成的反馈与针对性课本参考结合,应用于一门大型入门生物学课程。系统通过要求学生解释答案并声明信心水平,促进元认知行为;利用OpenAI的GPT-4o基于这些信息生成个性化反馈,并引导至相关课本章节。基于三场期中考试(参与人数分别为541、342、413人)的交互日志(共28,313次问答-学生互动,覆盖146个学习目标),以及279份问卷和23次访谈,我们分析了系统对学习成效与参与度的影响。所有反馈类型在成绩上无显著差异,但趋势显示潜在益处。最关键的成效来自强制的信心评分与解释要求,学生报告已将其迁移至真实考试策略中。约40%的学生在反馈提示下查阅了课本资料,远超传统阅读率。问卷显示满意度高(均值4.1/5),82.1%的学生表示对练习内容信心提升,73.4%称能回忆并应用具体概念。研究表明,嵌入结构化反思要求可能比复杂的反馈机制更具影响力。
原文摘要 · Abstract (English)
Providing personalized, detailed feedback at scale in large undergraduate STEM courses remains a persistent challenge. We present an empirically evaluated practice exam system that integrates AI generated feedback with targeted textbook references, deployed in a large introductory biology course. Our system encourages metacognitive behavior by asking students to explain their answers and declare their confidence. It uses OpenAI's GPT-4o to generate personalized feedback based on this information, while directing them to relevant textbook sections. Through interaction logs from consenting participants across three midterms (541, 342, and 413 students respectively), totaling 28,313 question-student interactions across 146 learning objectives, along with 279 surveys and 23 interviews, we examined the system's impact on learning outcomes and engagement. Across all midterms, feedback types showed no statistically significant performance differences, though some trends suggested potential benefits. The most substantial impact came from the required confidence ratings and explanations, which students reported transferring to their actual exam strategies. About 40 percent of students engaged with textbook references when prompted by feedback -- far higher than traditional reading rates. Survey data revealed high satisfaction (mean rating 4.1 of 5), with 82.1 percent reporting increased confidence on practiced midterm topics, and 73.4 percent indicating they could recall and apply specific concepts. Our findings suggest that embedding structured reflection requirements may be more impactful than sophisticated feedback mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。