针对新学生冷启动问题,用元学习快速适应少量初始答题数据。
MAML-KT: Addressing Cold Start Problem in Knowledge Tracing for New Students via Few-Shot Model-Agnostic Meta Learning

- 采用元学习思想,从少量数据快速适配新学生知识状态。
- 在3-10题窗口期,准确率显著高于DKT、SAKT等模型。
- 适合需要快速评估新学习者能力的个性化教学系统。
知识追踪(KT)模型通常在所有学生早期交互数据上训练,后期响应上测试。然而,这种评估方式掩盖了实际部署中常见的冷启动场景:模型需仅凭少数初始答题记录推断从未见过的学生知识状态。已有研究显示,标准经验风险最小化的KT模型(如DKT、DKVMN、SAKT)在此条件下早期准确率大幅下降。本文将新学生预测建模为少样本学习问题,提出MAML-KT——一种模型无关的元学习方法,通过优化初始化实现仅一次或两次梯度更新即可快速适应新学生。在ASSIST2009、ASSIST2015和ASSIST2017数据集上,采用控制性冷启动协议(训练子集学生,测试保留学习者在第3-10题与第11-15题窗口),并调整班级规模(10至50人)。结果表明,几乎所有冷启动条件下MAML-KT均取得更高早期准确率,且随班级规模增加仍保持优势。在ASSIST2017中,早期性能短暂下降,恰逢大量学生首次接触新技能。分析表明该下降源于技能新颖性而非模型不稳定,与先前关于技能级冷启动的研究一致。优化模型快速适应能力可降低新学生早期预测误差,并更清晰揭示早期准确率波动,区分模型缺陷与真实学习动态。
原文摘要 · Abstract (English)
Knowledge tracing (KT) models are commonly evaluated by training on early interactions from all students and testing on later responses. While effective for measuring average predictive performance, this evaluation design obscures a cold start scenario that arises in deployment, where models must infer the knowledge state of previously unseen students from only a few initial interactions. Prior studies have shown that under this setting, standard empirically risk-minimized KT models such as DKT, DKVMN and SAKT exhibit substantially lower early accuracy than previously reported. We frame new-student performance prediction as a few-shot learning problem and introduce MAML-KT, a model-agnostic meta learning approach that learns an initialization optimized for rapid adaptation to new students using one or two gradient updates. We evaluate MAML-KT on ASSIST2009, ASSIST2015 and ASSIST2017 using a controlled cold start protocol that trains on a subset of students and tests on held-out learners across early interaction windows (questions 3-10 and 11-15), scaling cohort sizes from 10 to 50 students. Across datasets, MAML-KT achieves higher early accuracy than prior KT models in nearly all cold start conditions, with gains persisting as cohort size increases. On ASSIST2017, we observe a transient drop in early performance that coincides with many students encountering previously unseen skills. Further analysis suggests that these drops coincide with skill novelty rather than model instability, consistent with prior work on skill-level cold start. Overall, optimizing KT models for rapid adaptation reduces early prediction error for new students and provides a clearer lens for interpreting early accuracy fluctuations, distinguishing model limitations from genuine learning and knowledge acquisition dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。