构建统一生存基准,揭示学习行为才是预测辍学的关键。
A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics

- 按动态周频与静态早期窗口分组对比多种生存模型
- 行为时间信号比人口统计更关键,尤其随机生存森林表现最优
- 适合教育数据科学家和学习分析研究者参考
学生辍学是学习分析中的长期难题,但现有研究常在异构评估协议下比较模型,侧重区分能力而忽视时间可解释性与校准性。本研究基于开放大学学习分析数据集(OULAD),构建面向时间辍学风险建模的生存基准。比较两类方法:家庭A(动态每周)采用个体-时期表示;家庭B(静态早期窗口)涵盖树模型、参数模型与神经网络。评估整合四个层面:预测性能、消融实验、可解释性与校准性。结果按家庭分别报告,避免因时间表征差异导致的误判。家庭B中,随机生存森林在所有三个时间点上均取得最高时变一致性指数与最低Brier分数;家庭A中,泊松分段指数模型在集成Brier分数上表现最佳,且五模型间差距极小。无重拟合自助抽样支持其为方向性信号,非严格优劣。消融与可解释性分析一致指出:主导预测信号并非人口或结构特征,而是时间与行为特征。校准结果验证此模式,仅XGBoost AFT为异常值。研究支持多维统一基准,并将辍学风险定位为时间-行为过程,而非静态背景属性函数。
原文摘要 · Abstract (English)
Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration. This study introduces a survival-oriented benchmark for temporal dropout risk modelling using the Open University Learning Analytics Dataset (OULAD). Two arms are compared: Family A: Dynamic Weekly, with models in person-period representation, and Family B: Static Early-Window, with an expanded roster of families: tree-based survival, parametric, and neural models. The evaluation protocol integrates four analytical layers: predictive performance, ablation, explainability, and calibration. Results are reported within each family separately, because a single numerical cross-family ranking would conflate genuine model differences with artifacts of temporal representation, to which survival metrics are known to be sensitive. Within Family B, Random Survival Forest showed the highest point estimates for time-dependent concordance and the lowest Brier scores across all three horizons; within Family A, Poisson Piecewise-Exponential showed the lowest point estimate for integrated Brier score within a tight five-model cluster. No-refit bootstrap resampling qualifies these positions as directional signals, not claims of strict superiority. Ablation and explainability analyses converged, across all models, on a shared finding: the dominant predictive signal was not primarily demographic or structural, but temporal and behavioral. Calibration corroborated this pattern in the better-discriminating models, except for XGBoost AFT, the sole outlier (analyzed in the Discussion). These results support unified, multi-dimensional benchmarking in Learning Analytics and situate dropout risk as a temporal-behavioral process rather than a function of static background attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。