LLMs常误判学生易错题,因它们只看教学顺序而非认知困难。
The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
- LLMs根据教学顺序评估难度,忽略学生常见误解。
- 分数题被高估简单度,如100÷1/2仅34.16%正确。
- 适合教育评估与自适应系统设计者警惕算法偏差。
大型语言模型(LLMs)在教育测评中用于估计题目难度,但其评估是否反映学习者真实体验尚不明确。本研究考察了基于4种主流LLM系统的难度评分(1-100分制)与770名印尼大学生实证表现之间的对齐程度。通过经典测验理论(CTT)和双参数项目反应理论(2PL)计算实际难度,共生成640次评分。结果显示,斯皮尔曼等级相关系数为0.52–0.70,表明模型能捕捉粗略的难度排序。然而,在分数运算题上出现显著系统性偏差:多个被模型判定为简单的题目,实际学生正确率极低,如100 ÷ 1/2仅有34.16%答对。我们认为,LLMs反映的是课程编排难度,即“应如何简单”,而非由认知误解驱动的真实难度。这种系统性低估现象被称为‘易陷阱’。研究揭示了基于LLM的难度估计的关键局限,提示缺乏实证验证的评估设计可能引入偏见。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。