融合情感分析与行为数据,提升在线学习辍学预测准确率
SentiDrop: A Multi Modal Machine Learning model for Predicting Dropout in Distance Learning
- 用BERT分析学生评论情感,结合XGBoost处理行为与人口统计特征
- 在下学年未见数据上达84%准确率,优于基线模型的82%
- 适合教育机构开发个性化干预策略,降低在线学习辍学率
在线学习中的辍学问题严重,早期预警对干预和学生坚持至关重要。利用教育数据预测辍学是学习分析领域的研究热点。我们合作的在线学习平台强调整合多种数据源的重要性,包括社会人口统计、行为数据和情感分析。本文提出一种新模型,将基于BERT的学生评论情感分析结果,与通过极端梯度提升(XGBoost)分析的社会人口统计及行为数据相结合。我们对BERT在学生评论上进行微调以捕捉细微情感,并使用特征重要性技术筛选关键特征后融合。模型在下一学年的未见数据上测试,准确率达84%,高于基线模型的82%。此外,模型在精确率和F1分数等指标上也表现更优。该方法可作为制定个性化策略以降低辍学率的重要工具。
原文摘要 · Abstract (English)
School dropout is a serious problem in distance learning, where early detection is crucial for effective intervention and student perseverance. Predicting student dropout using available educational data is a widely researched topic in learning analytics. Our partner's distance learning platform highlights the importance of integrating diverse data sources, including socio-demographic data, behavioral data, and sentiment analysis, to accurately predict dropout risks. In this paper, we introduce a novel model that combines sentiment analysis of student comments using the Bidirectional Encoder Representations from Transformers (BERT) model with socio-demographic and behavioral data analyzed through Extreme Gradient Boosting (XGBoost). We fine-tuned BERT on student comments to capture nuanced sentiments, which were then merged with key features selected using feature importance techniques in XGBoost. Our model was tested on unseen data from the next academic year, achieving an accuracy of 84\%, compared to 82\% for the baseline model. Additionally, the model demonstrated superior performance in other metrics, such as precision and F1-score. The proposed method could be a vital tool in developing personalized strategies to reduce dropout rates and encourage student perseverance
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。