arXiv:2509.25961cs.CL2025-09EMNLP被引 1

发现无参考纠错评估方法可被攻击,导致评分失真。

Reliability Crisis of Reference-free Metrics for Grammatical Error Correction

  • 设计对抗攻击策略,让纠错系统骗过无参考评估指标
  • 四种主流无参考指标均被攻破,攻击系统性能超越当前最优
  • 提醒需构建更可靠的自动评估方法,尤其对自动生成系统

无参考语法纠错评估指标虽与人工判断高度相关,但未针对旨在获得不合理高分的对抗性系统进行设计。此类系统存在会破坏自动评估的可靠性,误导用户选择错误的GEC系统。本研究针对SOME、Scribendi、IMPARA及基于LLM的四种无参考指标,提出对抗攻击策略,并验证所构建的对抗系统性能优于当前最先进水平。结果凸显了发展更鲁棒评估方法的紧迫性。

原文摘要 · Abstract (English)

Reference-free evaluation metrics for grammatical error correction (GEC) have achieved high correlation with human judgments. However, these metrics are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores. The existence of such systems undermines the reliability of automatic evaluation, as it can mislead users in selecting appropriate GEC systems. In this study, we propose adversarial attack strategies for four reference-free metrics: SOME, Scribendi, IMPARA, and LLM-based metrics, and demonstrate that our adversarial systems outperform the current state-of-the-art. These findings highlight the need for more robust evaluation methods.

语法纠错评估漏洞对抗攻击无参考评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。