arXiv:2508.03294cs.CLcs.AI2025-08被引 2

用大模型不确定性预测考题难度,比教授更准

NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty

  • 用大模型解题时的不确定度作为特征,做监督学习
  • 仅需42个样本即超越教授,准确率显著提升
  • 适合教育评估者、考试命题人快速测题难易

评估考试题目难度对制定优质试卷至关重要,但教授在该任务上表现有限。我们对比多种基于大语言模型的方法与三位教授在神经网络与机器学习领域判断真假题正确率的能力。结果表明,教授难以区分题目难易,且表现逊于直接让Gemini 2.5完成此任务。进一步地,在监督学习框架下使用大模型解题的不确定性,仅需42个训练样本便获得更优结果。结论:利用大模型不确定性进行监督学习,可有效辅助教授更准确估计试题难度,提升评估质量。

原文摘要 · Abstract (English)

Estimating the difficulty of exam questions is essential for developing good exams, but professors are not always good at this task. We compare various Large Language Model-based methods with three professors in their ability to estimate what percentage of students will give correct answers on True/False exam questions in the areas of Neural Networks and Machine Learning. Our results show that the professors have limited ability to distinguish between easy and difficult questions and that they are outperformed by directly asking Gemini 2.5 to solve this task. Yet, we obtained even better results using uncertainties of the LLMs solving the questions in a supervised learning setting, using only 42 training samples. We conclude that supervised learning using LLM uncertainty can help professors better estimate the difficulty of exam questions, improving the quality of assessment.

难度预测大模型教育评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。