arXiv:2509.21514cs.LGcs.CL2025-09

让知识追踪模型学会在不确定时主动放弃预测,提升可靠性。

Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

  • 用蒙特卡洛丢弃法量化模型不确定性,实现无需重训练的择优预测。
  • 放弃最不确定的20%预测,准确率提升2.3至3.0个百分点。
  • 揭示模型自身不确定性远超传统心理测量理论,适合需高可信度的教育场景。

知识追踪(KT)模型通常关注提升预测准确性,但负责任的实际应用要求模型能识别何时应将不确定预测交由人类教师处理。本文为现有KT模型引入基于蒙特卡洛丢弃法(MC-Dropout)的内在选择性预测层,以量化不确定性。在Eedi数学数据集上评估三种架构(DKT、SAKT、AKT),对最不确定的20%预测进行放弃,可使准确率提升2.3至3.0个百分点,AUC提升1.9至2.4个百分点,F1提升1.4至4.3个百分点,且无需重新训练。被放弃样本的错误率是保留样本的1.45至1.60倍,该策略在各题目难度四分位及学生能力水平中均保持公平。此外,MC-Dropout方差带来的AUC提升约为校准双参数逻辑模型(2PL IRT)基线的五倍。对模型认知不确定性(BALD)的分解显示,传统心理测量体系(包括题目难度、学生能力、IRT型结果模糊性及历史课程覆盖)在线性建模下仅解释不足4%的信号,非线性回归最多解释23%,剩余77%至90%为模型架构特有的认知不确定性,普通代理方法无法捕捉。因此,基于模型原生认知不确定性的选择性预测,是负责任部署知识追踪的必要组成部分,应与子群体公平性审计和课堂评估并行,而非替代。

原文摘要 · Abstract (English)

Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment requires models to know when to defer uncertain predictions to a human teacher. We introduce an intrinsic selective prediction layer for existing KT models using Monte Carlo Dropout (MC-Dropout) to quantify uncertainty. We evaluate this approach across three architectures (DKT, SAKT, and AKT) using the Eedi mathematics dataset. Abstaining on the 20\% most uncertain predictions lifts accuracy by 2.3 to 3.0 percentage points, AUC by 1.9 to 2.4 percentage points and F1 by 1.4 to 4.3 percentage points without any retraining. This abstention strategy is highly targeted: the deferred set exhibits 1.45 to 1.60 times the error rate of the kept set. Furthermore, this targeting holds within every question-difficulty quartile and remains fair across student-ability levels. Importantly, MC-Dropout variance gives roughly five times the AUC lift of a calibrated two-parameter logistic (2PL) Item Response Theory (IRT) baseline as a selective-prediction signal. A variance decomposition of the model's epistemic uncertainty (BALD) reveals that the entire classical psychometric stack, comprising question difficulty, student ability, IRT-style outcome ambiguity, and historical curriculum coverage, explains less than 4\% of the signal under linear modeling and at most 23\% even with a non-linear regressor. This leaves 77\% to 90\% as architecture-specific epistemic content that MC-Dropout surfaces and simpler proxies cannot recover. Selective prediction with model-native epistemic uncertainty is therefore a necessary component of responsible KT deployment, complementary to subgroup-fairness audits and downstream classroom evaluation rather than a substitute for them.

知识追踪不确定性教育AI选择性预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。