用AI评分填补缺失分数,提升答题测试能力评估精度
Leveraging AI Graders for Missing Score Imputation to Achieve Accurate Ability Estimation in Constructed-Response Tests
- 利用AI评分模型预测未人工批改的题目得分
- 在缺失率达50%时仍保持高准确度的能力估计
- 适合教育评估系统、大规模考试平台快速部署
学习者能力评估是教育领域的重要目标,尤其对表达能力和逻辑思维等高阶能力的考察需求日益增长。以简答题和论文题为代表的构造响应测试被广泛采用,但其依赖大量人工评分,成本高昂且效率低。项目反应理论(IRT)可通过不完整评分数据估计能力,但随着缺失分数比例上升,估计准确性显著下降。现有数据增强方法在稀疏或异构数据上表现不佳。本文提出一种新方法:利用自动化评分技术填补缺失分数,实现基于IRT的精准能力估计。该方法在显著降低人工评卷工作量的同时,保持了高水平的能力评估准确性。
原文摘要 · Abstract (English)
Evaluating the abilities of learners is a fundamental objective in the field of education. In particular, there is an increasing need to assess higher-order abilities such as expressive skills and logical thinking. Constructed-response tests such as short-answer and essay-based questions have become widely used as a method to meet this demand. Although these tests are effective, they require substantial manual grading, making them both labor-intensive and costly. Item response theory (IRT) provides a promising solution by enabling the estimation of ability from incomplete score data, where human raters grade only a subset of answers provided by learners across multiple test items. However, the accuracy of ability estimation declines as the proportion of missing scores increases. Although data augmentation techniques for imputing missing scores have been explored in order to address this limitation, they often struggle with inaccuracy for sparse or heterogeneous data. To overcome these challenges, this study proposes a novel method for imputing missing scores by leveraging automated scoring technologies for accurate IRT-based ability estimation. The proposed method achieves high accuracy in ability estimation while markedly reducing manual grading workload.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。