arXiv:2606.18709cs.CL2026-06被引 2

测试大模型能否区分不同水平学生的阅读题,发现效果有限。

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

论文配图:LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
图 1 · 摘自论文原文
  • 用直接预测和模拟答题两种方法评估模型对题目区分度的判断能力
  • 模型最高仅达到0.231的相关性,远未可靠捕捉学生差异
  • 当前模型无法稳定模拟不同能力水平学生的答题行为

现有基于大语言模型(LLM)的教育评估研究主要关注题目难度,但难度本身无法反映题目是否能有效区分高、低能力学生。题目区分度是衡量这一能力的互补且基础的心理测量属性。本文研究了大语言模型能否从题目内容中预测人类的题目区分度。我们采用两种互补方法评估42个私有及开源大模型:直接区分度预测要求模型显式预测题目的区分度值;响应型代理估计则将模型输出视为合成作答,并基于经典测验理论(CTT)思想进行题项-剩余项分析。直接预测与人类判断的匹配度较弱。响应型代理虽提供更强的排序信号,但在CEFR分层下相关性仅为0.231。进一步分析表明,该相关性主要来自模型间的差异,而非通过能力提示可靠模拟不同能力水平的学生表现。因此,当前大模型虽包含部分区分度相关信息,但尚未能可靠建模赋予题目区分度心理测量意义的能力依赖型人类作答行为。

原文摘要 · Abstract (English)

Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item meaningfully distinguishes higher- from lower-proficiency students. Item discrimination captures this complementary and fundamental psychometric property. We investigate whether LLMs can predict human item discrimination from assessment content. We evaluate 42 proprietary and open-weight LLMs using two complementary approaches. Direct discrimination prediction asks models to explicitly predict an item's discrimination value, while response-based proxy estimation treats LLM answers as synthetic responses and applies a Classical Test Theory (CTT)-inspired item-rest calculation. Direct predictions show weak alignment with human item discrimination. The response-based proxy provides a stronger but still limited ranking signal, reaching a CEFR-stratified rank correlation of 0.231. Further analysis shows that this correlation comes mainly from differences across models rather than proficiency prompts that reliably simulate students at different ability levels. Current LLMs therefore contain some discrimination-relevant information, but they do not yet reliably model the ability-conditioned human response behavior that gives item discrimination its psychometric meaning.

教育评估大模型心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。