开源大模型在托福作文评分中存在母语偏见,欧洲背景作文得分普遍更高。
Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

- 用微调的Gemma-3模型在12100篇跨语言作文上测试评分一致性。
- 整体评分准确率达77.79%,跨提示泛化能力强,无主题偏好。
- 发现欧洲语言背景作文平均得分高于东亚背景,且与训练数据无关。
本研究考察了基于LoRA微调的开源大语言模型(Gemma-3-27B-it)在自动作文评分中的跨提示泛化能力与母语(L1)评分偏差。采用与《AiAWE》(Gayed, 2026)相同的模型和推理配置,在480篇论说文上微调后,评估其在完整TOEFL11语料库上的表现:12,100篇来自11种母语背景的考生作文,涵盖8个未在训练中出现的提示。模型原始分数(0.5–5.0)映射至ETS使用的三个水平等级(低、中、高),实现直接对比。模型整体等级一致率达77.79%,加权卡帕系数为0.702,邻近等级一致率高达99.98%。在所有8个未见提示上表现稳定,无主题相关提示优势,表明强跨提示泛化能力。但模型存在系统性母语关联评分偏移:在每个等级内,欧洲语言背景作文得分均显著高于东亚语言背景,该现象无法由微调数据构成解释。这是首个针对微调开源大模型在自动作文评分中的大规模母语公平性分析。
原文摘要 · Abstract (English)
This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models" (Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。