arXiv:2502.09416cs.CL2025-02ACL被引 9

改进语法纠错评估方法,让机器评分更贴近人类判断。

Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?

  • 用人类对比打分的方式聚合自动评分结果
  • 多数指标在SEEDA上排名准确率提升
  • 连GPT-4的评分也不如改进后的方法

语法错误纠正(GEC)的自动评估目标是使系统排序与人类偏好一致。然而当前自动评估采用平均句级绝对得分的方式,与人类通过句子间相对比较并聚合打分的方法不同。本文提出一种与人类评价方式对齐的自动评估聚合方法,实验涵盖基于编辑、n-gram和句级的多种指标,在SEEDA基准上验证了该方法能显著提升多数指标的表现。结果显示,即使基于BERT的指标有时也优于GPT-4的评分。该方法已集成至gec-metrics工具库。

原文摘要 · Abstract (English)

One of the goals of automatic evaluation metrics in grammatical error correction (GEC) is to rank GEC systems such that it matches human preferences. However, current automatic evaluations are based on procedures that diverge from human evaluation. Specifically, human evaluation derives rankings by aggregating sentence-level relative evaluation results, e.g., pairwise comparisons, using a rating algorithm, whereas automatic evaluation averages sentence-level absolute scores to obtain corpus-level scores, which are then sorted to determine rankings. In this study, we propose an aggregation method for existing automatic evaluation metrics which aligns with human evaluation methods to bridge this gap. We conducted experiments using various metrics, including edit-based metrics, n-gram based metrics, and sentence-level metrics, and show that resolving the gap improves results for the most of metrics on the SEEDA benchmark. We also found that even BERT-based metrics sometimes outperform the metrics of GPT-4. The proposed ranking method is integrated gec-metrics.

语法纠错评估指标机器评价人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。