arXiv:2603.00608cs.AI2026-03

用机器学习同时预测学生挂科和成绩,提前干预辍学风险。

Machine Learning Grade Prediction Using Students' Grades and Demographics

  • 统一框架同步做分类和回归,比分开处理更全面。
  • 分类准确率达96%,回归模型决定系数达0.70。
  • 适合教育管理者用于早期识别高危学生群体。

中学阶段学生重修带来巨大资源压力,尤其在资源有限的地区。本研究提出一种统一的机器学习框架,同时预测通过/不及格结果与连续成绩,突破以往将分类与回归分开处理的局限。评估了六种模型:逻辑回归、决策树、随机森林用于分类;线性回归、决策树回归器、随机森林回归器用于回归,所有超参数通过穷举网格搜索优化。基于4424名中学生的学业与人口统计学数据,分类模型最高准确率达96%,回归模型决定系数(R²)达0.70,优于基线方法。结果证实了早期数据驱动识别高风险学生可行性,并凸显双任务联合预测的价值。该框架可支持及时个性化干预,有助于降低重修率并优化资源配置。

原文摘要 · Abstract (English)

Student repetition in secondary education imposes significant resource burdens, particularly in resource-constrained contexts. Addressing this challenge, this study introduces a unified machine learning framework that simultaneously predicts pass/fail outcomes and continuous grades, a departure from prior research that treats classification and regression as separate tasks. Six models were evaluated: Logistic Regression, Decision Tree, and Random Forest for classification, and Linear Regression, Decision Tree Regressor, and Random Forest Regressor for regression, with hyperparameters optimized via exhaustive grid search. Using academic and demographic data from 4424 secondary school students, classification models achieved accuracies of up to 96%, while regression models attained a coefficient of determination of 0.70, surpassing baseline approaches. These results confirm the feasibility of early, data-driven identification of at-risk students and highlight the value of integrating dual-task prediction for more comprehensive insights. By enabling timely, personalized interventions, the framework offers a practical pathway to reducing grade repetition and optimizing resource allocation.

教育预测机器学习学生画像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。