用真人校对耗时评估语法纠错工具效果,更贴近真实使用体验。
Time Is Effort: Estimating Human Post-Editing Time for Grammar Error Correction Tool Evaluation
- 基于真人校对时间构建首个大规模标注数据集
- 发现句子是否需要修改、改写和标点调整最耗时
- 新指标PEET与人工判断高度相关,适合评估工具实用性
文本编辑常需多次修订。在初稿修正阶段引入高效的语法纠错(GEC)工具,会显著影响后续人工编辑的投入与最终文本质量。这引发一个关键问题:GEC工具能为用户节省多少工作量?我们首次构建了针对两个英文GEC测试集(BEA19和CoNLL14)的大规模后编辑(PE)时间标注与修正数据集。提出以校对耗时为核心的后编辑努力度(PEET)指标,用于量化任何GEC工具的校对时间成本。利用该数据集,我们测算了不同工具带来的实际时间节省。分析表明,判断句子是否需修改,以及改写和标点调整类操作对校对时间影响最大。与人工评分对比显示,PEET与技术努力感知高度一致,为评估GEC工具可用性提供了全新的以人为核心的方向。数据与代码已开源:https://github.com/ankitvad/PEET_Scorer。
原文摘要 · Abstract (English)
Text editing can involve several iterations of revision. Incorporating an efficient Grammar Error Correction (GEC) tool in the initial correction round can significantly impact further human editing effort and final text quality. This raises an interesting question to quantify GEC Tool usability: How much effort can the GEC Tool save users? We present the first large-scale dataset of post-editing (PE) time annotations and corrections for two English GEC test datasets (BEA19 and CoNLL14). We introduce Post-Editing Effort in Time (PEET) for GEC Tools as a human-focused evaluation scorer to rank any GEC Tool by estimating PE time-to-correct. Using our dataset, we quantify the amount of time saved by GEC Tools in text editing. Analyzing the edit type indicated that determining whether a sentence needs correction and edits like paraphrasing and punctuation changes had the greatest impact on PE time. Finally, comparison with human rankings shows that PEET correlates well with technical effort judgment, providing a new human-centric direction for evaluating GEC tool usability. We release our dataset and code at: https://github.com/ankitvad/PEET_Scorer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。