用模拟学生表现评估题目质量,让自动出题更符合教学需求。
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
- 通过构建高质量与低质量题目对,模拟学生答题表现来评估题目
- 现有方法无法准确区分题目在难度、区分度等方面的差异
- 适合教育AI研究者和智能题库开发者参考
尽管问题生成(QG)任务在教育测评中日益普及,其评估仍受限于缺乏与测试题教育价值明确关联的方法。本文将教师常用的教学评价工具——题目分析法引入QG评估,构建在主题覆盖、题目难度、区分度和干扰项效率等维度上存在差异的题目对,检验现有评估方法是否能有效识别这些差异。实验发现,现有方法在关联学生表现方面存在明显不足。为此,我们提出QG-SMS框架,利用大语言模型进行学生建模与仿真,实现基于学生行为的题目分析。大量实验与人工评估表明,模拟学生画像带来的新视角显著提升了题目评估的有效性与鲁棒性。
原文摘要 · Abstract (English)
While the Question Generation (QG) task has been increasingly adopted in educational assessments, its evaluation remains limited by approaches that lack a clear connection to the educational values of test items. In this work, we introduce test item analysis, a method frequently used by educators to assess test question quality, into QG evaluation. Specifically, we construct pairs of candidate questions that differ in quality across dimensions such as topic coverage, item difficulty, item discrimination, and distractor efficiency. We then examine whether existing QG evaluation approaches can effectively distinguish these differences. Our findings reveal significant shortcomings in these approaches with respect to accurately assessing test item quality in relation to student performance. To address this gap, we propose a novel QG evaluation framework, QG-SMS, which leverages Large Language Model for Student Modeling and Simulation to perform test item analysis. As demonstrated in our extensive experiments and human evaluation study, the additional perspectives introduced by the simulated student profiles lead to a more effective and robust assessment of test items.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。