大模型难估学生难题,越强反而越不准。
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
- 用模拟不同水平学生的答题表现来测试模型难度感知能力。
- 模型越大越偏离人类难度判断,出现机器共识偏差。
- 模型无法识别自身局限,适合教育评估场景的改进研究。
准确估计题目难度对教育评估至关重要,但存在冷启动问题。尽管大语言模型展现出超人级解题能力,但其是否能感知学习者的认知困难仍不明确。本文针对超过20个模型在医学知识、数学推理等多元领域进行了大规模实证分析,发现模型与人类在难度判断上存在系统性错位:模型规模扩大并未提升对人类困难的感知,反而趋向统一的机器共识。高性能模型往往更难以准确估计难度,即使被明确要求模拟特定能力水平的学生表现,仍无法有效还原学习者的能力局限。此外,模型普遍缺乏自我反思能力,无法预测自身局限。结果表明,通用解题能力并不等于对人类认知困难的理解,提示当前模型在自动化难度预测中面临根本挑战。
原文摘要 · Abstract (English)
Accurate estimation of item (question or task) difficulty is critical for educational assessment but suffers from the cold start problem. While Large Language Models demonstrate superhuman problem-solving capabilities, it remains an open question whether they can perceive the cognitive struggles of human learners. In this work, we present a large-scale empirical analysis of Human-AI Difficulty Alignment for over 20 models across diverse domains such as medical knowledge and mathematical reasoning. Our findings reveal a systematic misalignment where scaling up model size is not reliably helpful; instead of aligning with humans, models converge toward a shared machine consensus. We observe that high performance often impedes accurate difficulty estimation, as models struggle to simulate the capability limitations of students even when being explicitly prompted to adopt specific proficiency levels. Furthermore, we identify a critical lack of introspection, as models fail to predict their own limitations. These results suggest that general problem-solving capability does not imply an understanding of human cognitive struggles, highlighting the challenge of using current models for automated difficulty prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。