用学生错误理解生成更贴近真实考试的多选题
Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
- 基于学生回答分析错误认知,生成包含错项的多选题
- 生成题目在难度和区分度上更接近人工命题
- 适合教育AI、智能评测系统开发者使用
本研究提出一种名为AnaQuest的创新提示技术,利用预训练大语言模型生成多选题(MCQs)。在该方法中,选项为关于复杂概念的句子级陈述。形式化评估阶段,学生以自由文本回答目标概念的问题;总结性评估阶段,AnaQuest分析这些回答,自动生成正确与错误的陈述。通过项目反应理论(IRT)对比分析,发现由AnaQuest生成的题目,尤其是错误选项(干扰项),在难度和区分度上比基线ChatGPT生成的题目更接近人工命题。专家评估也显示,两种AI生成的题目有效性与人类教师相当。
原文摘要 · Abstract (English)
The primary goal of this study is to develop and evaluate an innovative prompting technique, AnaQuest, for generating multiple-choice questions (MCQs) using a pre-trained large language model. In AnaQuest, the choice items are sentence-level assertions about complex concepts. The technique integrates formative and summative assessments. In the formative phase, students answer open-ended questions for target concepts in free text. For summative assessment, AnaQuest analyzes these responses to generate both correct and incorrect assertions. To evaluate the validity of the generated MCQs, Item Response Theory (IRT) was applied to compare item characteristics between MCQs generated by AnaQuest, a baseline ChatGPT prompt, and human-crafted items. An empirical study found that expert instructors rated MCQs generated by both AI models to be as valid as those created by human instructors. However, IRT-based analysis revealed that AnaQuest-generated questions - particularly those with incorrect assertions (foils) - more closely resembled human-crafted items in terms of difficulty and discrimination than those produced by ChatGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。