用大模型评估奥地利德语作文,准确率仍不理想。
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
- 用四大开源大模型按评分标准打分
- 最高仅40.6%子维度匹配人工评分
- 小模型无法胜任真实考试评分
自动作文评分(AES)旨在减轻教师负担并减少主观偏差,已有数十年研究。早期系统依赖手工特征与统计模型,而大语言模型(LLMs)的进展使写作评估更具灵活性。本文研究了最新开源大模型在奥地利A-level德语作文评分中的应用,聚焦基于评分量规的评估。使用101份匿名学生试卷,涵盖三种文本类型,评估了DeepSeek-R1 32b、Qwen3 30b、Mixtral 8x7b和LLama3.3 70b四种模型,采用不同上下文与提示策略。模型在量规子维度上最高达成40.6%的人工评分一致性,最终成绩仅有32.8%匹配人工专家评分。结果表明,尽管小模型能运用标准化量规,但准确度不足以投入实际评分环境。
原文摘要 · Abstract (English)
Automated Essay Scoring (AES) has been explored for decades with the goal to support teachers by reducing grading workload and mitigating subjective biases. While early systems relied on handcrafted features and statistical models, recent advances in Large Language Models (LLMs) have made it possible to evaluate student writing with unprecedented flexibility. This paper investigates the application of state-of-the-art open-weight LLMs for the grading of Austrian A-level German texts, with a particular focus on rubric-based evaluation. A dataset of 101 anonymised student exams across three text types was processed and evaluated. Four LLMs, DeepSeek-R1 32b, Qwen3 30b, Mixtral 8x7b and LLama3.3 70b, were evaluated with different contexts and prompting strategies. The LLMs were able to reach a maximum of 40.6% agreement with the human rater in the rubric-provided sub-dimensions, and only 32.8% of final grades matched the ones given by a human expert. The results indicate that even though smaller models are able to use standardised rubrics for German essay grading, they are not accurate enough to be used in a real-world grading environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。