arXiv:2602.00070cs.CYcs.CL2026-02被引 4

首个支持LLM教育研究的完整学生学习数据集,含真实答题与错因分析。

FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs

  • 提供完整题干、学生作答及错误选项选择记录,支持自然语言推理。
  • 170万条交互数据验证:模型在知识追踪上仅达基线水平,错题诊断能力低于随机。
  • 适合研究个性化教学、认知诊断与评估题目设计的教育AI团队使用。

大型语言模型能否理解学生的学习过程?随着大模型被用于自适应测试与个性化辅导,这一问题日益紧迫,但现有教育资源难以回答。当前教育数据集仅提供题目标识与对错标签,对基于自然语言推理的模型不透明。为此,我们推出FoundationalASSIST,首个英语教育数据集,包含完整题干、学生实际作答、错误选项选择记录,并对齐美国共同核心标准(Common Core K-12)。该数据集涵盖5,000名学生的170万条互动记录,支持此前无法开展的研究方向,如学生建模微调与误解模式分析。为验证其价值,我们在四个前沿模型(GPT-OSS-120B、Llama-3.3-70B、Qwen3-Next-80B变体)上评估了两类任务:知识追踪(预测学生答题表现与具体答案)与教学根基性评估(判断题目有效性特征)。结果揭示显著差距:所有模型在知识追踪上仅达基线水平;在题目区分度上均低于随机水平,表明模型未理解题目诊断性差异。尽管在相对难度判断上最高达68.6%,但仍凸显其他能力严重不足。这些发现表明,大模型尚需重大改进才能实现规模化个性化学习支持。我们公开发布FoundationalASSIST以推动基础性突破。

原文摘要 · Abstract (English)

Can Large Language Models understand how students learn? As LLMs are deployed for adaptive testing and personalized tutoring, this question becomes urgent -- yet we cannot answer it with existing resources. Current educational datasets provide only question identifiers and binary correctness labels, rendering them opaque to LLMs that reason in natural language. We address this gap with FoundationalASSIST, the first English educational dataset providing the complete information needed for research on LLMs in education: full question text, actual student responses (not just right/wrong), records of which wrong answers students chose, and alignment to Common Core K-12 standards. These 1.7 million interactions from 5,000 students enable research directions that were previously impossible to pursue, from fine-tuning student models to analyzing misconception patterns. To demonstrate the dataset's utility, we evaluate four frontier models (GPT-OSS-120B, Llama-3.3-70B, Qwen3-Next-80B variants) on two complementary task families: Knowledge Tracing, testing whether LLMs can predict student performance on questions, and the exact answer a student will give; and \textbf{Pedagogical Grounding}, testing whether LLMs understand the properties that make assessment items effective. Our evaluation reveals significant gaps in current LLM capabilities. Every model barely achieves a trivial baseline on knowledge tracing. All models fall below random chance on item discrimination, indicating that LLMs do not understand what makes one problem more diagnostic than another. Models do show competence at judging relative difficulty (up to 68.6%), but this partial success only highlights the gaps elsewhere. These results establish that substantial advances are needed before LLMs can reliably support personalized learning at scale. We release FoundationalASSIST to support progress on these foundational challenges.

教育AI知识追踪大模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。