用大模型和树模型预测中小学题目难度,减少试测成本。
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
- 用大模型提取认知语言特征,再用随机森林等模型预测难度
- 特征法相关系数达0.87,误差低于直接预测和传统方法
- 适合教育测评开发人员快速评估题目难易度
通过实地测试估算题目难度通常耗时费力。因此,仅基于题目内容大规模预测难度具有强烈需求。大型语言模型(LLMs)为此提供了新可能。本研究检验了使用LLM预测K-5年级数学与阅读测评题难度(共5170题)的可行性。采用两种方法:(a) 直接提示法,让LLM对每道题给出单一难度评分;(b) 特征法,由LLM提取多项认知与语言特征,再输入集成树模型(随机森林与梯度提升)进行预测。总体而言,直接LLM估计与真实难度呈中到强相关,但各年级表现不一,低年级常较差。相比之下,特征法预测精度更高,相关系数高达r = 0.87,误差低于直接预测与基线回归模型。结果表明,LLM有助于简化题目开发流程,降低对大规模试测的依赖,强调结构化特征提取的重要性。我们提供七步操作流程,供测评人员在其题库中实施类似方法。
原文摘要 · Abstract (English)
Estimating item difficulty through field-testing is often resource-intensive and time-consuming. As such, there is strong motivation to develop methods that can predict item difficulty at scale using only the item content. Large Language Models (LLMs) represent a new frontier for this goal. The present research examines the feasibility of using an LLM to predict item difficulty for K-5 mathematics and reading assessment items (N = 5170). Two estimation approaches were implemented: (a) a direct estimation method that prompted the LLM to assign a single difficulty rating to each item, and (b) a feature-based strategy where the LLM extracted multiple cognitive and linguistic features, which were then used in ensemble tree-based models (random forests and gradient boosting) to predict difficulty. Overall, direct LLM estimates showed moderate to strong correlations with true item difficulties. However, their accuracy varied by grade level, often performing worse for early grades. In contrast, the feature-based method yielded stronger predictive accuracy, with correlations as high as r = 0.87 and lower error estimates compared to both direct LLM predictions and baseline regressors. These findings highlight the promise of LLMs in streamlining item development and reducing reliance on extensive field testing and underscore the importance of structured feature extraction. We provide a seven-step workflow for testing professionals who would want to implement a similar item difficulty estimation approach with their item pool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。