用大模型不确定性预测选择题难度,效果优于现有方法。
Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation
- 通过大模型回答题目时的不确定性判断题目难易度。
- 在USMLE和CMCQRD数据集上达到最新最好结果。
- 适合教育科技、智能测评系统开发者参考。
在教育场景中,对多选题(MCQ)难度的估计对师生均有重要价值。由于人工评估成本高,自动化的题目难度估计受到关注,但此前效果参差不齐。本文提出新思路:让多个大语言模型解答三组不同MCQ数据集中的题目,利用模型输出的不确定性来预测题目难度。结合模型不确定性特征与文本特征,在随机森林回归器中验证发现,不确定性特征显著提升预测性能,且难度与正确作答学生比例成反比。实验表明,该方法在公开的USMLE和CMCQRD数据集上达到当前最优表现。
原文摘要 · Abstract (English)
In an educational setting, an estimate of the difficulty of multiple-choice questions (MCQs), a commonly used strategy to assess learning progress, constitutes very useful information for both teachers and students. Since human assessment is costly from multiple points of view, automatic approaches to MCQ item difficulty estimation are investigated, yielding however mixed success until now. Our approach to this problem takes a different angle from previous work: asking various Large Language Models to tackle the questions included in three different MCQ datasets, we leverage model uncertainty to estimate item difficulty. By using both model uncertainty features as well as textual features in a Random Forest regressor, we show that uncertainty features contribute substantially to difficulty prediction, where difficulty is inversely proportional to the number of students who can correctly answer a question. In addition to showing the value of our approach, we also observe that our model achieves state-of-the-art results on the USMLE and CMCQRD publicly available datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。