整合行为数据与认知校准分析,预测学生成绩并揭示其自我评估偏差。
An Integrated Machine Learning and Hierarchical Variance Decomposition Pipeline for Student Performance Prediction and Metacognitive Calibration on Multi-Signal Telemetry

- 构建三模块集成框架,融合成绩预测、校准误差评估与方差分解
- 学生自我评估误差显著高于模型,表明存在系统性认知偏差
- 可帮助教育AI系统个性化调整反馈,适合教育数据分析研究者
预测学生成绩与刻画元认知校准对智能辅导系统的个性化至关重要。以往研究将成绩预测、校准误差计算与方差分解视为独立流程,难以统一解释。本文提出统一行为预测与校准分析管道(UBP-CAP),通过三个联动模块处理学生执行前的行为遥测数据:(1) 使用LightGBM结合SHAP的二分类正确性预测;(2) 引入正式校准指标(ECE、MCE、Brier分数分解)评估元认知一致性;(3) 采用交叉广义线性混合效应模型(GLMM)分解校准偏差。提出预测-解释分歧指数(PEDI),量化预测与解释特征谱之间的结构性差异。在1,195条交互记录(27名学生,45个任务)上评估,逻辑回归达到AUC-ROC = 0.903,优于LightGBM(0.878)。学生原始ECE(0.109)显著高于模型ECE(0.068),证实系统性校准偏差。交叉GLMM显示ICCStudent = 0.123,表明校准具有情境性而非特质性。PEDIcos = 0.081(p = 0.327)说明预测与解释在共享行为特征上结构一致。
原文摘要 · Abstract (English)
Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems. Prior research treats performance prediction, calibration error calculation, and variance decomposition as separate pipelines, preventing unified interpretation. I propose the Unified Behavioral Prediction and Calibration Analysis Pipeline (UBP-CAP), an integrated framework processing student pre-execution behavioral telemetry through three linked modules: (1) a LightGBM classifier with SHAP for binary correctness prediction, (2) formal calibration metrics (ECE, MCE, and Brier score decomposition) to evaluate metacognitive alignment, and (3) a crossed Generalized Linear Mixed-Effects Model (GLMM) for decomposing calibration deviations. I introduce the Predictive-Explanatory Divergence Index (PEDI), which quantifies structural divergence between predictive and explanatory feature profiles. Evaluated on 1,195 interaction records (27 students, 45 tasks), Logistic Regression achieves AUC-ROC = 0.903, outperforming LightGBM (0.878). Student naive ECE (0.109) significantly exceeds model ECE (0.068), confirming systematic miscalibration. The crossed GLMM yields ICCStudent = 0.123, showing calibration is situational rather than dispositional. PEDIcos = 0.081 (p = 0.327) indicates structural alignment between prediction and explanation on shared behavioral features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。