提出新方法缓解自适应测试中的题目选择偏差,提升诊断公平性。
Selective Mixup for Debiasing Question Selection in Computerized Adaptive Testing
- 用均衡样本作为参考,通过选择性mixup增强偏差冲突样本多样性。
- 在两个基准数据集上显著提升诊断模型的泛化能力与公平性。
- 适合关注教育AI公平性、自适应测试优化的研究者和开发者。
计算机自适应测试(CAT)广泛用于在线教育平台评估学习者能力。通过基于先前能力估计选择题目并根据回答迭代更新估计,实现个性化建模并受到广泛关注。然而,现有工作多聚焦诊断准确率,忽视了自适应过程中的选择偏差。该偏差源于题目选择依赖于能力估计,导致低能力者被分配较易题目,高能力者被分配较难题目,进而使偏差在迭代中传播并放大,造成预测失准。此外,学习者历史交互数据的不平衡性进一步加剧了这一问题。为此,本文提出一种去偏框架,包含两个关键模块:跨属性考生检索与选择性mixup正则化。首先,检索具有正确/错误回答均衡分布的平衡考生作为偏差考生的中性参考;随后,在标签一致条件下,对每个偏差考生与其匹配的平衡样本进行mixup操作。该增强方式丰富了偏差冲突样本,平滑了题目选择边界。大量实验在两个基准数据集上使用多种先进诊断模型验证,结果表明该方法显著提升了CAT中题目选择的泛化能力与公平性。
原文摘要 · Abstract (English)
Computerized Adaptive Testing (CAT) is a widely used technology for evaluating learners' proficiency in online education platforms. By leveraging prior estimates of proficiency to select questions and updating the estimates iteratively based on responses, CAT enables personalized learner modeling and has attracted substantial attention. Despite this progress, most existing works focus primarily on improving diagnostic accuracy, while overlooking the selection bias inherent in the adaptive process. Selection Bias arises because the question selection is strongly influenced by the estimated proficiency, such as assigning easier questions to learners with lower proficiency and harder ones to learners with higher proficiency. Since the selection depends on prior estimation, this bias propagates into the diagnosis model, which is further amplified during iterative updates, leading to misalignment and biased predictions. Moreover, the imbalanced nature of learners' historical interactions often exacerbates the bias in diagnosis models. To address this issue, we propose a debiasing framework consisting of two key modules: Cross-Attribute Examinee Retrieval and Selective Mixup-based Regularization. First, we retrieve balanced examinees with relatively even distributions of correct and incorrect responses and use them as neutral references for biased examinees. Then, mixup is applied between each biased examinee and its matched balanced counterpart under label consistency. This augmentation enriches the diversity of bias-conflicting samples and smooths selection boundaries. Finally, extensive experiments on two benchmark datasets with multiple advanced diagnosis models demonstrate that our method substantially improves both the generalization ability and fairness of question selection in CAT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。