arXiv:2605.07635cs.CL2026-05被引 1

评测大模型语法纠错能力,发现最佳表现者并揭示评分漏洞

Multi-Dimensional Evaluation of LLMs for Grammatical Error Correction

  • 多维评估最新大模型在纠错精度、语流保持和语义保留上的表现
  • GPT-4o 纠错效果最优,且94.7%错误修正模式与其他模型高度一致
  • 73.76%的纠错结果虽不符标准答案,但质量相当或更优,提示现有评分方式严重低估系统性能

当前语法纠错自动化助手已广泛应用于服务数百万学习者的教育平台,但该领域仍存在三大关键空白:(1)最新一代大语言模型(LLMs)缺乏对语法纠错任务的全面评估;(2)多个模型融合是否提升纠错质量尚无研究;(3)基于参考文本的评估指标对纠错系统性能的低估程度未被充分量化。本研究首先在编辑精度、语言流畅性与语义保留三个维度上评估最新大模型,发现微调后的GPT-4o在三项指标上均达到领先水平。其次,通过语法错误类型分析表明,各模型的纠错模式高度相似(ρ=0.947)。第三,结果显示73.76%的GPT-4o纠错结果虽不同于人工标准答案,但质量和表现相当甚至更优。这些发现为教育工作者选择真正促进学生语言发展的纠错工具提供依据。相关数据、代码与模型均已公开。

原文摘要 · Abstract (English)

Automated assistants for Grammatical Error Correction are now embedded in educational platforms serving millions of learners, yet three critical gaps remain in this domain: (1) latest-generation Large Language Models (LLMs) lack comprehensive evaluation on grammar correction tasks; (2) whether combining these LLMs improves correction quality is unexplored; and (3) the extent to which reference-based metrics underestimate GEC system performance has not been adequately quantified. In this study, first, we evaluate latest-generation LLMs on edit precision, fluency preservation, and meaning retention, showing fine-tuned GPT-4o achieves state-of-the-art performance across all three dimensions. Second, through grammatical error type analysis we demonstrate that individual LLMs exhibit highly similar error correction patterns ($ρ=0.947$). Third, we show that reference-based metrics underestimate GEC performance with 73.76% of GPT-4o corrections different from gold standards being equally valid or even superior. These GEC evaluation findings equip educators with guidance for selecting GEC assistants that enhance rather than constrain student linguistic development. We make our data, code, and models publicly available.

语法纠错大模型评估GPT-4o教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。