用大模型分析学生长期学习数据,提前预测学业走向
Leveraging Language Models for Analyzing Longitudinal Experiential Data in Education
- 用预训练语言模型融合多源学习行为数据,增强时间模式学习能力
- 模型对缺失数据有韧性,能跨模态整合信息但依赖统计规律
- 适合教育数据科学与早期干预研究者参考
我们提出一种新方法,利用预训练语言模型(LMs)基于高维纵向体验数据,对理工科学生进行学业轨迹的早期预测。该数据涵盖学生的学习活动、行为及心理状态,有助于开展基于预测的干预。核心挑战包括高缺失率、因收集成本高导致的数据集规模有限,以及跨模态的时间复杂性。我们的方法通过全面的数据增强流程,结合缺失值处理、数据扩充,并嵌入任务特定指令与上下文提示,提升模型对时间模式的学习能力。在精心构建的学生学习数据集上,我们评估了编码器-解码器和仅解码器两类LMs。实验表明,尽管模型能有效跨模态整合数据并具备抗缺失能力,但仍主要依赖高层统计模式,对时间动态的深层理解不足,显式时间信息的解析能力有限。本研究推动了教育数据科学的发展,揭示了语言模型在基于纵向体验数据建模学生轨迹以实现早期干预中的潜力与局限。
原文摘要 · Abstract (English)
We propose a novel approach to leveraging pre-trained language models (LMs) for early forecasting of academic trajectories in STEM students using high-dimensional longitudinal experiential data. This data, which captures students' study-related activities, behaviors, and psychological states, offers valuable insights for forecasting-based interventions. Key challenges in handling such data include high rates of missing values, limited dataset size due to costly data collection, and complex temporal variability across modalities. Our approach addresses these issues through a comprehensive data enrichment process, integrating strategies for managing missing values, augmenting data, and embedding task-specific instructions and contextual cues to enhance the models' capacity for learning temporal patterns. Through extensive experiments on a curated student learning dataset, we evaluate both encoder-decoder and decoder-only LMs. While our findings show that LMs effectively integrate data across modalities and exhibit resilience to missing data, they primarily rely on high-level statistical patterns rather than demonstrating a deeper understanding of temporal dynamics. Furthermore, their ability to interpret explicit temporal information remains limited. This work advances educational data science by highlighting both the potential and limitations of LMs in modeling student trajectories for early intervention based on longitudinal experiential data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。