arXiv:2410.01480stat.MLcs.LG2024-10被引 2

提出新型多项选择题模型,提升测试分析精度并引入可比的比特评分尺度。

Introducing Flexible Monotone Multiple Choice Item Response Theory Models and Bit Scales

  • 用自编码器拟合单调多项选择模型,更贴合数据分布。
  • 在真实考试数据中,新模型拟合度优于传统名义反应模型。
  • 提出比特尺度,便于不同模型间分数比较,尤其适合无分布假设模型。

项目反应理论(IRT)是一种通过响应分析评估试题与被试能力的强大统计方法。拟合度更高的模型能带来更准确的潜在特质估计。本文提出一种针对多项选择数据的新模型——单调多项选择(MMC)模型,并使用自编码器进行拟合。通过模拟场景和瑞典学业能力测验的真实数据,实证表明,MMC模型在拟合度上优于传统的名义反应IRT模型。此外,本文展示了如何将任意已拟合的IRT模型的潜在特质尺度转换为比率尺度,有助于分数解释,并使不同类型IRT模型间的比较更加简便。这种新尺度称为比特尺度(bit scales),特别适用于对潜在特质分布假设极少或没有假设的模型,如本研究中的自编码器拟合模型。

原文摘要 · Abstract (English)

Item Response Theory (IRT) is a powerful statistical approach for evaluating test items and determining test taker abilities through response analysis. An IRT model that better fits the data leads to more accurate latent trait estimates. In this study, we present a new model for multiple choice data, the monotone multiple choice (MMC) model, which we fit using autoencoders. Using both simulated scenarios and real data from the Swedish Scholastic Aptitude Test, we demonstrate empirically that the MMC model outperforms the traditional nominal response IRT model in terms of fit. Furthermore, we illustrate how the latent trait scale from any fitted IRT model can be transformed into a ratio scale, aiding in score interpretation and making it easier to compare different types of IRT models. We refer to these new scales as bit scales. Bit scales are especially useful for models for which minimal or no assumptions are made for the latent trait scale distributions, such as for the autoencoder fitted models in this study.

项目反应理论自编码器比特尺度多选题建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。