用可解释AI分析10国数学成绩,找出影响因素差异。
Explainable AI for Predicting and Understanding Mathematics Achievement: A Cross-National Analysis of PISA 2018
- 用随机森林等模型分析6.7万学生数据,结合解释技术提升可读性。
- 非线性模型准确率更高,家庭背景和学习时间是关键预测因子。
- 揭示不同国家影响因素差异,适合教育政策制定者参考。
理解影响学生数学表现的因素对制定有效教育政策至关重要。本研究采用可解释人工智能(XAI)技术分析PISA 2018数据,预测数学成就并识别十个国家(67,329名学生)的关键预测因子。测试了四种模型:多元线性回归(MLR)、随机森林(RF)、CATBoost和人工神经网络(ANN),使用学生、家庭与学校变量。模型在70%数据上训练(5折交叉验证),30%数据上测试,按国家分层。性能以R²和平均绝对误差(MAE)评估。为保证可解释性,采用特征重要性、SHAP值和决策树可视化。非线性模型,尤其是随机森林和神经网络,优于线性模型,其中随机森林在准确率与泛化能力间取得平衡。关键预测因子包括社会经济地位、学习时间、教师动机及学生对数学的态度,但其影响因国家而异。散点图显示随机森林与CATBoost的预测值与实际得分高度一致。研究揭示成就的非线性与情境依赖性,凸显XAI在教育研究中的价值,有助于发现跨国规律、推动公平改革,并支持个性化学习策略开发。
原文摘要 · Abstract (English)
Understanding the factors that shape students' mathematics performance is vital for designing effective educational policies. This study applies explainable artificial intelligence (XAI) techniques to PISA 2018 data to predict math achievement and identify key predictors across ten countries (67,329 students). We tested four models: Multiple Linear Regression (MLR), Random Forest (RF), CATBoost, and Artificial Neural Networks (ANN), using student, family, and school variables. Models were trained on 70% of the data (with 5-fold cross-validation) and tested on 30%, stratified by country. Performance was assessed with R^2 and Mean Absolute Error (MAE). To ensure interpretability, we used feature importance, SHAP values, and decision tree visualizations. Non-linear models, especially RF and ANN, outperformed MLR, with RF balancing accuracy and generalizability. Key predictors included socio-economic status, study time, teacher motivation, and students' attitudes toward mathematics, though their impact varied across countries. Visual diagnostics such as scatterplots of predicted vs actual scores showed RF and CATBoost aligned closely with actual performance. Findings highlight the non-linear and context-dependent nature of achievement and the value of XAI in educational research. This study uncovers cross-national patterns, informs equity-focused reforms, and supports the development of personalized learning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。