arXiv:2503.18982cs.LGcs.AI2025-03

用生成对抗网络填补学习者答题数据空缺,提升智能辅导系统评估精度。

Generative Data Imputation for Sparse Learner Performance Data Using Generative Adversarial Imputation Networks

  • 基于GAN构建三维数据框架,灵活处理不同缺失率的学习行为数据。
  • 在多个真实教育数据集上,比张量分解和传统GAN方法更准确填补缺失值。
  • 填补后的数据能可靠还原学习行为模式,适合个性化教学系统使用。

智能辅导系统收集的学习者答题数据对知识状态建模至关重要,但因跳过或未完成导致的数据稀疏性影响评估与个性化教学。为此,我们提出基于生成对抗性数据填补网络(GAIN)的生成式填补方法。该方法采用学习者、题目和作答次数的三维框架,可适应不同缺失水平。通过卷积神经网络增强并采用最小二乘损失优化,使输入输出维度对齐于学习者-作答矩阵。在AutoTutor成人阅读理解(ARC)、ASSISTments和MATHia数据集上的实验表明,该方法在不同作答场景下显著优于张量分解及其它GAN方法的填补准确率。贝叶斯知识追踪(BKT)进一步验证:填补数据能有效估计初始掌握概率(P(L0))、学习率(P(T))、猜测率(P(G))和失误率(P(S)),模型拟合度更高,分布接近原始数据。KL散度分析显示偏差极小,证明填补数据有效保留了学习本质特征。结果表明,GAIN在缓解数据稀疏问题方面表现稳健,支持自适应个性化教学,提升评估精度与教育效果。

原文摘要 · Abstract (English)

Learner performance data collected by Intelligent Tutoring Systems (ITSs), such as responses to questions, is essential for modeling and predicting learners' knowledge states. However, missing responses due to skips or incomplete attempts create data sparsity, challenging accurate assessment and personalized instruction. To address this, we propose a generative imputation approach using Generative Adversarial Imputation Networks (GAIN). Our method features a three-dimensional (3D) framework (learners, questions, and attempts), flexibly accommodating various sparsity levels. Enhanced by convolutional neural networks and optimized with a least squares loss function, the GAIN-based method aligns input and output dimensions to question-attempt matrices along the learners' dimension. Extensive experiments using datasets from AutoTutor Adult Reading Comprehension (ARC), ASSISTments, and MATHia demonstrate that our approach significantly outperforms tensor factorization and alternative GAN methods in imputation accuracy across different attempt scenarios. Bayesian Knowledge Tracing (BKT) further validates the effectiveness of the imputed data by estimating learning parameters: initial knowledge (P(L0)), learning rate (P(T)), guess rate (P(G)), and slip rate (P(S)). Results indicate the imputed data enhances model fit and closely mirrors original distributions, capturing underlying learning behaviors reliably. Kullback-Leibler (KL) divergence assessments confirm minimal divergence, showing the imputed data preserves essential learning characteristics effectively. These findings underscore GAIN's capability as a robust imputation tool in ITSs, alleviating data sparsity and supporting adaptive, individualized instruction, ultimately leading to more precise and responsive learner assessments and improved educational outcomes.

数据填补智能辅导GAN教育数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。