大模型能精准评估外语写作的多维度表现,反馈质量可靠。
LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing
- 用9项标准评估留学生论文,让大模型生成评分与评语。
- 大模型在多维评估中表现合理且结果可靠。
- 提供可复现的数据集和代码,适合教育研究者使用。
本文研究大模型在多维度分析性写作评估中的表现,即能否基于多个评估标准同时给出分数与评语。利用由二语研究生撰写的文献综述语料库(经人工专家按9项分析性标准评分),我们测试了多个主流大模型在不同提示条件下的评估能力。为评估反馈评语质量,提出一种新型可解释、低成本、可扩展且可复现的评价框架,相比依赖人工判断的现有方法更具优势。结果表明,大模型能够生成合理且总体可靠的多维度分析性评估。论文公开了数据集与代码,支持研究复现。
原文摘要 · Abstract (English)
The paper explores the performance of LLMs in the context of multi-dimensional analytic writing assessments, i.e. their ability to provide both scores and comments based on multiple assessment criteria. Using a corpus of literature reviews written by L2 graduate students and assessed by human experts against 9 analytic criteria, we prompt several popular LLMs to perform the same task under various conditions. To evaluate the quality of feedback comments, we apply a novel feedback comment quality evaluation framework. This framework is interpretable, cost-efficient, scalable, and reproducible, compared to existing methods that rely on manual judgments. We find that LLMs can generate reasonably good and generally reliable multi-dimensional analytic assessments. We release our corpus and code for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。