用联邦学习预测挂科风险,保护隐私且效果好
Evaluating Federated Learning for At-Risk Student Prediction: A Comparative Analysis of Model Complexity and Data Balancing
- 在多地数据不共享前提下,用联邦学习联合训练模型
- 模型预测准确率约85%(ROC AUC),能有效识别高风险学生
- 适合关注教育隐私与早期预警的机构或研究者
本研究提出并验证了一个联邦学习(FL)框架,用于在保护数据隐私的前提下,主动识别存在辍学风险的学生。远程教育中持续存在的高辍学率仍是机构面临的重大挑战。基于大规模OULAD数据集,我们模拟了以隐私为中心的场景,利用早期学业表现和数字参与度数据训练模型。研究探讨了模型复杂度(逻辑回归与深度神经网络)与本地数据平衡策略之间的实际权衡。最终构建的联邦模型展现出较强的预测能力(ROC AUC约85%),证明联邦学习是兼顾学生数据主权、可扩展且可行的早期预警解决方案。
原文摘要 · Abstract (English)
This study proposes and validates a Federated Learning (FL) framework to proactively identify at-risk students while preserving data privacy. Persistently high dropout rates in distance education remain a pressing institutional challenge. Using the large-scale OULAD dataset, we simulate a privacy-centric scenario where models are trained on early academic performance and digital engagement patterns. Our work investigates the practical trade-offs between model complexity (Logistic Regression vs. a Deep Neural Network) and the impact of local data balancing. The resulting federated model achieves strong predictive power (ROC AUC approximately 85%), demonstrating that FL is a practical and scalable solution for early-warning systems that inherently respects student data sovereignty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。