MCQ和开放题学习效果相当,但MCQ更省时。
Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
- 对比三种练习方式:仅选择题、仅开放题、两者结合
- 三组在测验成绩上无显著差异,但选择题组用时更少
- 用GPT-4自动评分开放题可行,适合低风险评估
多项选择题(MCQ)作为学习工具的有效性一直存在争议。尽管其评分便捷,但随着大语言模型(LLM)实现自动评分,开放回答题日益用于教学。本研究通过后测仅随机对照实验,评估了六节倡导课程中三种练习方式(仅MCQ、仅开放题、两者结合)的学习效果。共234名导师完成790次课程。结果显示,三组在后测表现上无显著差异,但仅使用MCQ的组别完成时间显著更短。研究进一步利用GPT-4o与GPT-4-turbo对开放题进行自动评分,表明模型在低风险评估中表现良好,但广泛适用仍需更多研究。研究还公开了课程日志数据、人工评分标准与LLM提示模板,以促进透明与可复现性。
原文摘要 · Abstract (English)
The role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。