专用知识追踪模型比大模型更准更快更省。
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs
- 用小而专的模型追踪学生答题表现,针对性更强。
- 在准确率和F1上显著优于大模型,差距达数个百分点。
- 适合教育平台实时预测,部署成本低,响应速度快。
预测学生对题目未来的回答对教育平台至关重要,可实现有效干预。知识追踪(KT)模型是针对特定教育领域、基于学生答题数据训练的小型时序模型,具有高精度、快速推理和可扩展部署的优势。随着大语言模型(LLMs)兴起,我们提出三个问题:LLMs在此任务上的表现如何?是否具备可扩展性?与KT模型相比如何?本文对比了多个LLMs与KT模型在预测性能、部署成本和推理速度上的表现。结果表明,KT模型在该特定任务上准确率和F1得分均显著优于LLMs;且推理速度比LLMs快数个数量级,部署成本也低数个数量级。这说明教育预测任务中,专用模型依然优于通用大模型,当前闭源大模型不适合作为万能解决方案。
原文摘要 · Abstract (English)
Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions. One of the key approaches to do this has been through the use of knowledge tracing (KT) models. These are small, domain-specific, temporal models trained on student question-response data. KT models are optimised for high accuracy on specific educational domains and have fast inference and scalable deployments. The rise of Large Language Models (LLMs) motivates us to ask the following questions: (1) How well can LLMs perform at predicting students' future responses to questions? (2) Are LLMs scalable for this domain? (3) How do LLMs compare to KT models on this domain-specific task? In this paper, we compare multiple LLMs and KT models across predictive performance, deployment cost, and inference speed to answer the above questions. We show that KT models outperform LLMs with respect to accuracy and F1 scores on this domain-specific task. Further, we demonstrate that LLMs are orders of magnitude slower than KT models and cost orders of magnitude more to deploy. This highlights the importance of domain-specific models for education prediction tasks and the fact that current closed source LLMs should not be used as a universal solution for all tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。