三所大学合作预测学生留校率,不交换数据也能建模。
A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction

- 用远程数据科学框架,分高低侧服务器实现无数据共享建模。
- 跨机构分类准确率稳定在宏平均F1 0.690~0.695之间。
- 提出新型合成数据模板,更注重隐私而非数据分布相似性。
本研究探索基于PySyft平台的隐私保护机器学习(PPML)技术,实现机构间学生留校率的协作预测。我们构建了一种半隔离架构的远程数据科学(RDS)框架,包含高侧与低侧服务器,使三所大学的研究人员能在不直接访问敏感学生数据的情况下建立预测模型。基于一所小型私立大学的历史数据(N=720),评估了三种合成数据生成方法,并通过跨机构协作验证了该框架的有效性。结果表明,在严格遵守《家庭教育权利与隐私法案》(FERPA)的前提下,各机构间分类性能一致(宏平均F1:0.690–0.695)。我们还提出了数据类型感知模板(Data-Type-Aware Templates),一种以隐私优先于分布保真度的合成数据生成方法。研究证实,基于RDS的PPML在教育场景中技术可行,为小规模跨机构合作提供了优于联邦学习的实际替代方案。代码已公开于https://github.com/jtfields/NAIRR240195-Privacy-Preserving-Machine-Learning。
原文摘要 · Abstract (English)
This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that prioritizes privacy over distributional fidelity. Our findings confirm that RDS-based PPML is technically feasible for educational settings and offers a practical alternative to federated learning for small-scale inter-institutional collaborations. The code is available at https://github.com/jtfields/NAIRR240195-Privacy-Preserving-Machine-Learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。