用可解释特征预测爱沙尼亚语学习者水平,准确率超90%
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
- 选特定语言特征训练模型,提升可解释性
- 最高准确率达0.9,跨文本类型表现稳定
- 适合语言评估系统开发者和教育研究者
利用自然语言处理分析真实学习者语言,有助于构建自动化评估与反馈工具,并揭示二语产出发展的新规律。然而,现有研究较少同时关注这两方面。本研究旨在分类爱沙尼亚语水平考试写作(A2-C1级),通过精心筛选特征,构建更可解释、泛化性更强的机器学习模型。分析训练数据的语言属性,识别与复杂度和正确性相关的熟练度预测因子,包括词汇、形态、表层及错误特征。使用预选特征训练分类模型,其测试准确率与包含其他特征的模型相当,但不同文本类型的分类波动更小。最佳分类器准确率约0.9。对7-10年前考试样本的额外评估显示,写作复杂度上升,部分特征集仍保持0.8准确率。结果已集成至爱沙尼亚开源语言学习平台的写作评估模块。
原文摘要 · Abstract (English)
Using NLP to analyze authentic learner language helps to build automated assessment and feedback tools. It also offers new and extensive insights into the development of second language production. However, there is a lack of research explicitly combining these aspects. This study aimed to classify Estonian proficiency examination writings (levels A2-C1), assuming that careful feature selection can lead to more explainable and generalizable machine learning models for language testing. Various linguistic properties of the training data were analyzed to identify relevant proficiency predictors associated with increasing complexity and correctness, rather than the writing task. Such lexical, morphological, surface, and error features were used to train classification models, which were compared to models that also allowed for other features. The pre-selected features yielded a similar test accuracy but reduced variation in the classification of different text types. The best classifiers achieved an accuracy of around 0.9. Additional evaluation on an earlier exam sample revealed that the writings have become more complex over a 7-10-year period, while accuracy still reached 0.8 with some feature sets. The results have been implemented in the writing evaluation module of an Estonian open-source language learning environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。