arXiv:2504.02916cs.CYcs.LG2025-04被引 2

从学习系统数据中提取关键特征,提升学生表现预测精度。

Feature Engineering on LMS Data to Optimize Student Performance Prediction

  • 基于登录与成绩数据构建多种特征工程方法
  • 发现疫情后数据模式变化,影响历史数据适用性
  • 为教育数据分析提供可复用的特征设计指南

几乎所有教育机构都使用学习管理系统(LMS),通常产生由数千人生成的海量数据。本文分析了一所区域性综合性大学的LMS成绩与登录数据,重点记录了在预测学生表现时从这些数据中进行特征工程的关键考量。特别指出新冠疫情以来LMS数据模式的变化,这对数据科学家使用历史数据具有重要影响。通过对比多种特征工程方法及其在机器学习中的应用效果,验证了不同特征组合的有效性。最后总结了将这些特征纳入更全面的学生表现预测模型的实践意义。

原文摘要 · Abstract (English)

Nearly every educational institution uses a learning management system (LMS), often producing terabytes of data generated by thousands of people. We examine LMS grade and login data from a regional comprehensive university, specifically documenting key considerations for engineering features from these data when trying to predict student performance. We specifically document changes to LMS data patterns since Covid-19, which are critical for data scientists to account for when using historic data. We compare numerous engineered features and approaches to utilizing those features for machine learning. We finish with a summary of the implications of including these features into more comprehensive student performance models.

特征工程教育数据学生预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。