arXiv:2509.16224cs.CYcs.CL2025-09被引 1

用动机陈述文本分析预测大一辍学,效果不输传统成绩数据

Predicting First Year Dropout from Pre Enrolment Motivation Statements Using Text Mining

  • 用TF-IDF、主题建模和LIWC词典分析入学动机文本
  • 仅靠文本分析就能达到与学生背景特征相当的预测准确率
  • 为早期干预提供新思路,适合教育数据挖掘研究者参考

防止大学生辍学是高等教育的重大挑战,但入学前难以预测哪些学生会辍学。高中平均绩点虽是强预测因子,但仍无法解释大部分辍学差异。本研究利用文本挖掘技术分析学生提交的入学动机陈述,旨在提取其中隐含的信息。通过将文本数据与传统辍学预测指标(如学生特征)结合,尝试扩展预测维度。数据集包含2014至2015年荷兰一所非选拔性本科院校7,060份动机陈述。使用支持向量机在75%的数据上训练模型,并在测试集上评估多个组合模型(包括TF-IDF、主题建模、LIWC词典)。结果显示,尽管文本与传统特征的组合未提升预测效果,但仅用文本分析即可达到与学生特征集合相当的辍学预测能力。研究提出未来可进一步探索文本特征的深层语义价值。

原文摘要 · Abstract (English)

Preventing student dropout is a major challenge in higher education and it is difficult to predict prior to enrolment which students are likely to drop out and which students are likely to succeed. High School GPA is a strong predictor of dropout, but much variance in dropout remains to be explained. This study focused on predicting university dropout by using text mining techniques with the aim of exhuming information contained in motivation statements written by students. By combining text data with classic predictors of dropout in the form of student characteristics, we attempt to enhance the available set of predictive student characteristics. Our dataset consisted of 7,060 motivation statements of students enrolling in a non-selective bachelor at a Dutch university in 2014 and 2015. Support Vector Machines were trained on 75 percent of the data and several models were estimated on the test data. We used various combinations of student characteristics and text, such as TFiDF, topic modelling, LIWC dictionary. Results showed that, although the combination of text and student characteristics did not improve the prediction of dropout, text analysis alone predicted dropout similarly well as a set of student characteristics. Suggestions for future research are provided.

辍学预测文本挖掘教育数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。