arXiv:2410.21033stat.MLcs.LG2024-10被引 7

用机器学习快速校准题库并自适应出题,提升考试效率与精准度。

BanditCAT and AutoIRT: Machine Learning Approaches to Computerized Adaptive Testing and Item Calibration

  • 结合AutoML与项目反应理论,用题干特征自动校准试题参数。
  • 通过贝叶斯更新实时估计考生能力,仅需少量作答即可精准评分。
  • 适合大规模在线考试系统,尤其适用于新题型快速上线场景。

本文提出一个完整的框架,可在少量答题数据下快速校准并实施稳健的大规模计算机化自适应测试(CAT)。校准环节采用AutoIRT方法,结合自动化机器学习(AutoML)与项目反应理论(IRT),先使用基于题干特征的非参数AutoML评分模型,再训练特定项目的参数化模型,从而构建可解释的IRT模型。本研究采用表格型AutoML工具(AutoGluon.tabular)、BERT嵌入及语言学动机的NLP特征。测试实施中采用BanditCAT框架,将问题建模为上下文关联的老虎机问题,以项目在给定考生能力下的费舍尔信息作为奖励信号,利用汤普森采样平衡探索与利用,实现高区分度题目的选择。为控制题目曝光率,在计算费舍尔信息前引入随机化噪声。该框架已用于两个新题型在DET练习测试中的首次发布,基于5次实验评估了其可靠性与题目暴露度指标。

原文摘要 · Abstract (English)

In this paper, we present a complete framework for quickly calibrating and administering a robust large-scale computerized adaptive test (CAT) with a small number of responses. Calibration - learning item parameters in a test - is done using AutoIRT, a new method that uses automated machine learning (AutoML) in combination with item response theory (IRT), originally proposed in [Sharpnack et al., 2024]. AutoIRT trains a non-parametric AutoML grading model using item features, followed by an item-specific parametric model, which results in an explanatory IRT model. In our work, we use tabular AutoML tools (AutoGluon.tabular, [Erickson et al., 2020]) along with BERT embeddings and linguistically motivated NLP features. In this framework, we use Bayesian updating to obtain test taker ability posterior distributions for administration and scoring. For administration of our adaptive test, we propose the BanditCAT framework, a methodology motivated by casting the problem in the contextual bandit framework and utilizing item response theory (IRT). The key insight lies in defining the bandit reward as the Fisher information for the selected item, given the latent test taker ability from IRT assumptions. We use Thompson sampling to balance between exploring items with different psychometric characteristics and selecting highly discriminative items that give more precise information about ability. To control item exposure, we inject noise through an additional randomization step before computing the Fisher information. This framework was used to initially launch two new item types on the DET practice test using limited training data. We outline some reliability and exposure metrics for the 5 practice test experiments that utilized this framework.

自适应测试机器学习题库校准贝叶斯更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。