构建首个针对中小学生英语写作的细粒度错误分析基准
FEANEL: A Benchmark for Fine-Grained Error Analysis in K-12 English Writing
- 基于词性分类法构建错误标注体系
- 1000篇学生作文含专家标注的错误类型与反馈
- 揭示大模型在教育反馈上的显著短板
大型语言模型(LLMs)在人工智能领域带来深远变革,为教育应用提供巨大机遇。然而,其在中小学英语写作中提供细粒度教育反馈的能力仍待深入探索。本文提出细粒度英语学习者错误分析问题,并构建了细粒度错误分析英语学习者(FEANEL)基准。该基准包含1000篇中小学生撰写的作文,以及一套由语言教育专家共同开发的英语写作错误分类体系。每个错误均按类型、严重程度和解释性反馈进行标注,采用基于词性的分类方法。我们对当前最先进的大模型在FEANEL基准上进行了评估,结果揭示了现有模型在细粒度错误分析和教学能力方面存在显著差距,凸显了教育应用场景下技术改进的迫切需求。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have transformed artificial intelligence, offering profound opportunities for educational applications. However, their ability to provide fine-grained educational feedback for K-12 English writing remains underexplored. In this paper, we challenge the error analysis and pedagogical skills of LLMs by introducing the problem of Fine-grained Error Analysis for English Learners and present the Fine-grained Error ANalysis for English Learners (FEANEL) Benchmark. The benchmark comprises 1,000 essays written by elementary and secondary school students, and a well-developed English writing error taxonomy. Each error is annotated by language education experts and categorized by type, severity, and explanatory feedback, using a part-of-speech-based taxonomy they co-developed. We evaluate state-of-the-art LLMs on the FEANEL Benchmark to explore their error analysis and pedagogical abilities. Experimental results reveal significant gaps in current LLMs' ability to perform fine-grained error analysis, highlighting the need for advancements in particular methods for educational applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。