arXiv:2507.08232cs.CLcs.AI2025-07中稿 · the 20th Workshop …被引 6

用真实学生数据测试大模型,发现它无法稳定模拟各年级学生水平。

Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?

  • 用项目反应理论将大模型与真实学生放在同一能力量表上对比
  • 无提示时强模型普遍高于平均学生水平,弱模型偶然对齐
  • 提示调整后仍无法跨学科跨年级准确模拟,需新训练策略

大型语言模型(LLMs)被广泛用作智能辅导系统和试题预测试的虚拟学生。然而,这些虚拟学生在多大程度上能真实模拟真实学生的表现仍不明确。为此,我们收集了来自国家教育进展评估(NAEP)的489道题目,涵盖小学四年级、八年级和十二年级的数学与阅读理解。通过项目反应理论(IRT)模型,我们将11种不同且先进的大模型置于与真实学生群体相同的认知能力量表上进行定位。结果表明,在无引导情况下,强通用模型在所有年级均显著优于平均学生水平;而较弱或领域不匹配的模型可能偶然对齐。使用年级强化提示虽能改变模型表现,但其是否与对应年级平均学生水平一致高度依赖模型与提示组合:在各科目与年级间均无任何模型-提示组合达到理想匹配,凸显了新型训练与评估策略的必要性。最后,我们基于研究结果提出虚拟学生代理的选择指南。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as proxy students in the development of Intelligent Tutoring Systems (ITSs) and in piloting test questions. However, to what extent these proxy students accurately emulate the behavior and characteristics of real students remains an open question. To investigate this, we collected a dataset of 489 items from the National Assessment of Educational Progress (NAEP), covering mathematics and reading comprehension in grades 4, 8, and 12. We then apply an Item Response Theory (IRT) model to position 11 diverse and state-of-the-art LLMs on the same ability scale as real student populations. Our findings reveal that, without guidance, strong general-purpose models consistently outperform the average student at every grade, while weaker or domain-mismatched models may align incidentally. Using grade-enforcement prompts changes models' performance, but whether they align with the average grade-level student remains highly model- and prompt-specific: no evaluated model-prompt pair fits the bill across subjects and grades, underscoring the need for new training and evaluation strategies. We conclude by providing guidelines for the selection of viable proxies based on our findings.

大模型评估教育测评虚拟学生能力量表

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。