用大模型模拟学生答题,低成本预测数学题真实难度
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
- 让大模型扮演不同年级学生,通过角色扮演模拟答题过程
- 预测难度与真实数据相关性达0.75~0.82,表现优于直接评分
- 用多样姓名和性别种族分层能提升预测效果,弱数学模型反而更优
标准化数学测评需耗费人力进行试测以确定题目难度。本文研究开源大语言模型(LLMs)在评估真实学生多选题难度方面的预测能力。尽管大模型直接判断难度表现不佳,但通过模拟学生角色的生成式方法,在特定条件下可取得良好效果。我们通过提示大模型扮演4、8或12年级不同水平的学生,构建虚拟‘课堂’,利用模拟结果拟合项目反应理论(IRT)模型,并与美国教育进展评估(NAEP)提供的真实题目难度数据对比。结果显示,各年级题目正确率预测相关系数分别达到0.75、0.76和0.82。实验还考察了不同‘班级规模’的影响,发现角色命名多样性(如使用真实姓名)及按性别、种族分层可提升预测性能。有趣的是,数学能力较弱的模型(Gemma)比更强的模型(Llama和Qwen)更能准确预测真实难度,表明该任务更适合非强数学能力模型。
原文摘要 · Abstract (English)
Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language models (LLMs) for evaluating the difficulty of multiple-choice math questions for real-world students. We show that, while LLMs are poor direct judges of problem difficulty, simulation-based approaches with LLMs yield promising results under the right conditions. Under the proposed approach, we simulate a ``classroom'' of 4th, 8th, or 12th-grade students by prompting the LLM to role-play students of varying proficiency levels. We use the outcomes of these simulations to fit Item Response Theory (IRT) models, comparing learned difficulty parameters for items to their real-world difficulties, as determined by item-level statistics furnished by the National Assessment of Educational Progress (NAEP). We observe correlations as high as 0.75, 0.76, and 0.82 for grades 4, 8, and 12, respectively, on the item-level correctness rates. In our simulations, we experiment on math MCQs with different ``classroom sizes,'' showing tradeoffs between computation size and accuracy. We find that role-plays with diverse-named students improve predictions (compared to student IDs), and stratifying names across gender and race further improves predictions. Our results show that LLMs with relatively weaker mathematical abilities (Gemma) actually yield better real-world difficulty predictions than mathematically stronger models (Llama and Qwen), further underscoring the suitability of these models for the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。