首个匈牙利语反思写作水平自动分类研究,解决教育评估难题。
Automatic Reflection Level Classification in Hungarian Student Essays

- 用TF-IDF与语义嵌入特征结合传统机器学习模型
- 在1954篇标注作文上实现71%综合性能(准确率/F1/ROC AUC)
- 为低资源语言和不均衡数据提供可复用的分析框架
反思性思维是教育中的关键能力,但人工评估反思写作耗时且主观。尽管已有多种语言的自动化分析研究,匈牙利语仍缺乏系统探索。本文首次对匈牙利语学生作文的反思水平进行自动分类研究,使用包含1,954篇经专家标注的反思作文、按四级量表标记的大规模数据集。我们比较两类方法:(1) 基于TF-IDF与语义嵌入特征的传统机器学习模型;(2) 针对文档级反思分类微调的匈牙利语专用Transformer模型。针对数据集严重类别不平衡问题,系统考察了类别权重、过采样、数据增强与替代损失函数。通过详尽消融实验分析各策略贡献。结果表明,经过适当特征工程的浅层模型达到71%综合性能(准确率、F1-score、ROC AUC均值),而基于Transformer的模型综合性能略低(68%),但在少数类上表现出更好泛化能力。研究揭示传统方法在低资源场景下的有效性,以及Transformer模型在不均衡分类中的鲁棒性。所提出的数据集与实验洞察为匈牙利语及其他形态丰富的语言的自动化反思分析研究奠定基础。
原文摘要 · Abstract (English)
Reflective thinking is a key competency in education, but assessing reflective writing remains a time-consuming and subjective task for education experts. While automated reflective analysis has been explored in several languages, Hungarian language was not researched extensively. In this paper, we present the first comprehensive study on automatic reflection level classification in Hungarian student essays. We used a large, expert-annotated Hungarian dataset consisting of 1,954 reflective essays collected over multiple academic years and labeled on a four-level reflection scale. We investigate two approaches: (1) classical machine learning models using TF-IDF and semantic embedding features, and (2) Hungarian-specific transformer models fine-tuned for document-level reflection classification. To address the strong class imbalance in the dataset, we systematically examine class weighting, oversampling, data augmentation, and alternative loss functions. An extensive ablation study is conducted to analyze the contribution of each modeling and balancing strategy. Our results show that shallow machine learning models with appropriate feature engineering achieve strong overall performance, reaching up to 71% overall score averaged over accuracy, F1-score, and ROC AUC metrics, while transformer-based models achieve slightly lower overall score (68%) averaged over the same metrics, but demonstrate better generalization on minority reflection classes. These findings highlight the continued relevance of classical methods for low-resource settings and the robustness of transformer models for imbalanced classification. The proposed dataset and experimental insights provide a solid foundation for future research on automated reflective analysis in Hungarian and other morphologically rich languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。