研究词级质量评估对人工校对效率与质量的影响
QE4PE: Word-level Quality Estimation for Human Post-Editing
- 对比四种错误标记方式,评估其在真实校对场景中的效果
- 发现领域、语言和校对速度显著影响标记有效性
- 自动化标记虽准但实用性不及人工,反映准确率与可用性差距
词级质量评估(QE)旨在识别机器翻译中的错误片段,以指导人工校对。尽管词级QE系统的准确性已被广泛评估,其在真实场景中对人工校对速度、质量及编辑决策的影响仍研究不足。本研究在42名专业校对员参与的两个翻译方向的真实工作流中,评估了四种错误片段标记方式——包括监督式与不确定性驱动的词级QE方法——在先进神经机器翻译模型输出上的表现。通过行为日志估算校对投入与效率,结合词级与句段级人工标注评估质量提升。结果表明,领域、语言及校对员速度是决定标记效果的关键因素,且人工标记与自动标记的效果差异微小,凸显了专业流程中准确性与可用性之间的鸿沟。
原文摘要 · Abstract (English)
Word-level quality estimation (QE) methods aim to detect erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the speed, quality and editing choices of human post-editing remain understudied. In this study, we investigate the impact of word-level QE on machine translation (MT) post-editing in a realistic setting involving 42 professional post-editors across two translation directions. We compare four error-span highlight modalities, including supervised and uncertainty-based word-level QE methods, for identifying potential errors in the outputs of a state-of-the-art neural MT model. Post-editing effort and productivity are estimated from behavioral logs, while quality improvements are assessed by word- and segment-level human annotation. We find that domain, language and editors' speed are critical factors in determining highlights' effectiveness, with modest differences between human-made and automated QE highlights underlining a gap between accuracy and usability in professional workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。