用AI预测考试题难易度,准确率超80%。
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
- 用语言模型自动分析题目文本,无需人工设计特征。
- 模型预测误差最低达0.165,相关性高达0.87。
- 适合教育评估、智能命题系统研究者参考。
题目的难度在大规模测评中对成绩表现、分数解释和公平性至关重要。传统方法依赖实地测试与经典测验理论(CTT)或项目反应理论(IRT)校准,耗时且成本高。为克服此问题,基于文本的机器学习与语言模型方法成为有前景的替代方案。本文系统综述了截至2025年5月发表的37篇自动化题难易度预测研究。每项研究均涵盖数据集、难度参数、学科领域、题型、题目数量、训练测试划分、输入、特征、模型、评估标准及性能结果。结果显示,尽管经典机器学习因可解释性仍具价值,但先进语言模型(含小规模与大规模Transformer架构)能有效捕捉句法与语义模式,无需人工特征工程。尤为关键的是,模型性能被总结为未来研究基准:最低均方根误差(RMSE)达0.165,皮尔逊相关系数最高达0.87,准确率最高达0.806。论文最后讨论实践意义并提出未来研究方向。
原文摘要 · Abstract (English)
Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on field testing and classical test theory (CTT)-based item analysis or item response theory (IRT) calibration, which can be time-consuming and costly. To overcome these challenges, text-based approaches leveraging machine learning and language models, have emerged as promising alternatives. This paper reviews and synthesizes 37 articles on automated item difficulty prediction in large-scale assessment settings published through May 2025. For each study, we delineate the dataset, difficulty parameter, subject domain, item type, number of items, training and test data split, input, features, model, evaluation criteria, and model performance outcomes. Results showed that although classic machine learning models remain relevant due to their interpretability, state-of-the-art language models, using both small and large transformer-based architectures, can capture syntactic and semantic patterns without the need for manual feature engineering. Uniquely, model performance outcomes were summarized to serve as a benchmark for future research and overall, text-based methods have the potential to predict item difficulty with root mean square error (RMSE) as low as 0.165, Pearson correlation as high as 0.87, and accuracy as high as 0.806. The review concludes by discussing implications for practice and outlining future research directions for automated item difficulty modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。