LLM做学生建模效果差,难替代传统学习者建模技术。
Problems With Large Language Models for Learner Modelling: Why LLMs Alone Fall Short for Responsible Tutoring in K--12 Education
- 对比DKT与大模型,用真实数据评估知识追踪准确性
- DKT预测正确率最高(AUC=0.83),远超微调后的大模型
- 大模型存在错误更新、不连贯轨迹,不适合责任型教育应用
大型语言模型(LLM)在基础教育中的快速应用导致了一个误解:生成式模型可取代传统学习者建模实现自适应教学。这一问题在欧盟《人工智能法案》定义为高风险的K--12教育场景中尤为严重。本研究综合现有证据并实证检验了关键缺陷:评估学生随时间演化的知识水平的准确度、可靠性与时序一致性。通过对比深度知识追踪(DKT)模型与广泛使用的大型语言模型(零样本与微调版本),使用大规模公开数据集进行评估。结果显示,DKT在预测下一步正确性上表现最佳(AUC = 0.83),且在所有测试场景中持续优于大模型。尽管微调使大模型的AUC比零样本提升约8%,但仍落后于DKT 6%,且早期序列错误更高,对自适应支持危害最大。时序分析进一步显示,DKT保持稳定且方向正确的掌握度更新,而大模型版本则存在显著时序缺陷,包括不一致和方向错误的更新。这些缺陷即使在微调模型耗时近198小时、计算成本远超DKT的情况下依然存在。定性分析表明,微调后大模型仍产生不一致的多技能掌握轨迹,而DKT保持平滑连贯的更新。总体而言,结果表明仅靠大模型难以达到成熟智能辅导系统的有效性,负责任的辅导需采用融合学习者建模的混合框架。
原文摘要 · Abstract (English)
The rapid rise of large language model (LLM)-based tutors in K--12 education has fostered a misconception that generative models can replace traditional learner modelling for adaptive instruction. This is especially problematic in K--12 settings, which the EU AI Act classifies as high-risk domain requiring responsible design. Motivated by these concerns, this study synthesises evidence on limitations of LLM-based tutors and empirically investigates one critical issue: the accuracy, reliability, and temporal coherence of assessing learners' evolving knowledge over time. We compare a deep knowledge tracing (DKT) model with a widely used LLM, evaluated zero-shot and fine-tuned, using a large open-access dataset. Results show that DKT achieves the highest discrimination performance (AUC = 0.83) on next-step correctness prediction and consistently outperforms the LLM across settings. Although fine-tuning improves the LLM's AUC by approximately 8\% over the zero-shot baseline, it remains 6\% below DKT and produces higher early-sequence errors, where incorrect predictions are most harmful for adaptive support. Temporal analyses further reveal that DKT maintains stable, directionally correct mastery updates, whereas LLM variants exhibit substantial temporal weaknesses, including inconsistent and wrong-direction updates. These limitations persist despite the fine-tuned LLM requiring nearly 198 hours of high-compute training, far exceeding the computational demands of DKT. Our qualitative analysis of multi-skill mastery estimation further shows that, even after fine-tuning, the LLM produced inconsistent mastery trajectories, while DKT maintained smooth and coherent updates. Overall, the findings suggest that LLMs alone are unlikely to match the effectiveness of established intelligent tutoring systems, and that responsible tutoring requires hybrid frameworks that incorporate learner modelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。