arXiv:2410.01727cs.LGcs.CL2024-10被引 14

用大模型自动标注知识点,提升学习追踪模型效果

Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing

  • 用大模型生成解题步骤并自动标注知识点
  • 通过对比学习生成语义丰富的题目嵌入,提升表征能力
  • 可适配现有学习追踪模型,适合教育AI研究者

知识追踪(KT)是建模学生学习进展的常用方法,有助于实现个性化和自适应学习。然而现有方法存在两大局限:(1)依赖人工定义题目中的知识点(KCs),耗时且易出错;(2)忽略题目与知识点的语义信息。本文提出KCQRL框架,实现知识点的自动化标注与题目表示学习,可提升任意现有KT模型的效果。首先,利用大语言模型生成解题过程,并在每一步中自动标注知识点;其次,设计对比学习方法生成题目与解题步骤的丰富语义嵌入,通过定制化的负样本消除策略将嵌入与对应知识点对齐。这些嵌入可直接替换现有模型的随机初始化嵌入。我们在两个大规模真实数学学习数据集上验证了该方法,对15种不同KT算法均实现一致性能提升。

原文摘要 · Abstract (English)

Knowledge tracing (KT) is a popular approach for modeling students' learning progress over time, which can enable more personalized and adaptive learning. However, existing KT approaches face two major limitations: (1) they rely heavily on expert-defined knowledge concepts (KCs) in questions, which is time-consuming and prone to errors; and (2) KT methods tend to overlook the semantics of both questions and the given KCs. In this work, we address these challenges and present KCQRL, a framework for automated knowledge concept annotation and question representation learning that can improve the effectiveness of any existing KT model. First, we propose an automated KC annotation process using large language models (LLMs), which generates question solutions and then annotates KCs in each solution step of the questions. Second, we introduce a contrastive learning approach to generate semantically rich embeddings for questions and solution steps, aligning them with their associated KCs via a tailored false negative elimination approach. These embeddings can be readily integrated into existing KT models, replacing their randomly initialized embeddings. We demonstrate the effectiveness of KCQRL across 15 KT algorithms on two large real-world Math learning datasets, where we achieve consistent performance improvements.

知识追踪大模型自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。