arXiv:2502.18632cs.AIcs.CL2025-02ACL被引 6

用大模型自动生成编程题的知识点标签,提升学习追踪精度。

Automated Knowledge Component Generation for Interpretable Knowledge Tracing in Coding Problems

  • 基于大模型自动提取编程题的知识点并打标。
  • 新方法在预测学生答题结果上优于传统方法和人工标注。
  • 生成的知识点更符合认知规律,适合教育平台使用。

知识点(KCs)与题目映射有助于建模学生学习过程,跟踪其对细粒度技能的掌握程度,从而支持在线学习平台中的个性化教学与反馈。然而,传统上由领域专家手工构建和标注知识点耗时耗力。本文提出一种基于大模型的自动化知识点生成与标注流程,用于开放性编程题,并构建了基于该流程的可解释知识追踪框架(KCGen-KT)。我们在两种不同编程语言的真实学生代码提交数据集上进行了广泛的定量与定性评估。结果表明,KCGen-KT在预测未来学生答题表现方面优于现有知识追踪方法及人工标注的知识点。我们还分析了生成知识点的学习曲线,发现其在认知模型下拟合效果优于人工标注。此外,课程教师的人工评估显示,该流程生成的问题-知识点映射具有合理准确性。

原文摘要 · Abstract (English)

Knowledge components (KCs) mapped to problems help model student learning, tracking their mastery levels on fine-grained skills thereby facilitating personalized learning and feedback in online learning platforms. However, crafting and tagging KCs to problems, traditionally performed by human domain experts, is highly labor intensive. We present an automated, LLM-based pipeline for KC generation and tagging for open-ended programming problems. We also develop an LLM-based knowledge tracing (KT) framework to leverage these LLM-generated KCs, which we refer to as KCGen-KT. We conduct extensive quantitative and qualitative evaluations on two real-world student code submission datasets in different programming languages.We find that KCGen-KT outperforms existing KT methods and human-written KCs on future student response prediction. We investigate the learning curves of generated KCs and show that LLM-generated KCs result in a better fit than human written KCs under a cognitive model. We also conduct a human evaluation with course instructors to show that our pipeline generates reasonably accurate problem-KC mappings.

知识追踪大模型编程教育自动化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。