用联邦学习与差分特征预测学业风险学生,保护隐私且效果接近中心化模型。
Ranking-Based At-Risk Student Prediction Using Federated Learning and Differential Features
- 联邦学习实现跨校数据协作不共享原始数据,保障隐私。
- 差分特征提升模型性能,各数据集上预测准确率均优于传统方法。
- 可早期预警,学期初即有效识别高风险学生,适合教学干预场景。
数字教材在高校课程和在线讲授中广泛应用,生成的学习日志数据被用于教育数据挖掘(EDM)中的学生行为分析与成绩预测。然而,因隐私顾虑,学校间难以整合包含学业记录和学习日志的敏感数据,导致研究常受限于单一学校数据,难以构建高性能且泛化能力强的模型。本文提出结合联邦学习与差分特征的方法以解决该问题。联邦学习使模型训练无需集中数据,有效保护学生隐私;差分特征通过使用相对值而非绝对值,提升模型性能与泛化能力。实验基于12门课程、4年跨度共1,136名学生的数据训练模型,并在另外5门课程的保留测试集上验证。结果表明,所提方法在保持隐私的前提下,达到与集中式学习相当的性能,各项指标如Top-n精度、nDCG和PR-AUC均表现良好。此外,采用差分特征的模型在所有测试数据集上均优于非差分方法。模型还具备早期预测能力,在验证数据集中可于学期初期即高效识别高风险学生。
原文摘要 · Abstract (English)
Digital textbooks are widely used in various educational contexts, such as university courses and online lectures. Such textbooks yield learning log data that have been used in numerous educational data mining (EDM) studies for student behavior analysis and performance prediction. However, these studies have faced challenges in integrating confidential data, such as academic records and learning logs, across schools due to privacy concerns. Consequently, analyses are often conducted with data limited to a single school, which makes developing high-performing and generalizable models difficult. This study proposes a method that combines federated learning and differential features to address these issues. Federated learning enables model training without centralizing data, thereby preserving student privacy. Differential features, which utilize relative values instead of absolute values, enhance model performance and generalizability. To evaluate the proposed method, a model for predicting at-risk students was trained using data from 1,136 students across 12 courses conducted over 4 years, and validated on hold-out test data from 5 other courses. Experimental results demonstrated that the proposed method addresses privacy concerns while achieving performance comparable to that of models trained via centralized learning in terms of Top-n precision, nDCG, and PR-AUC. Furthermore, using differential features improved prediction performance across all evaluation datasets compared to non-differential approaches. The trained models were also applicable for early prediction, achieving high performance in detecting at-risk students in earlier stages of the semester within the validation datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。