让机器翻译评估更懂术语变化,避免误罚人类译者常用的灵活表达。
Improving Term Evaluation in Machine Translation: Variation Matters
- 引入跨术语变异检测,衡量翻译是否保留源语言的变体关系。
- 发现机器翻译的术语变化比人类少,且不同变体类型影响转移效果。
- 建议按是否保留源端变体关系来调整一致性惩罚,更适合真实翻译场景。
机器翻译中的术语评估通常假设每个源术语仅对应一个正确目标形式,但人类译者常使用多种表达方式,现有指标因此将这些变化视为不一致而惩罚。本文研究英文-法文科学文本翻译中如何纳入术语变化因素,在文档级评估中结合基于术语表的准确率、翻译一致性及一种新的跨术语变异(CTV)诊断指标,检验不同语言间变体关系是否被保留。基于两组平行语料库和四种MT系统的分析显示:(1)机器翻译生成的目标侧变体少于人类;(2)变体转移模式高度依赖于变体类型;(3)一致性排名随评估指标不同而变化;(4)使用术语表约束虽提升准确率与一致性,却削弱了CTV,抑制了合理变体。因此主张应建立考虑变体关系的评估机制,仅在目标侧变体反映源侧变体时才施加一致性惩罚。
原文摘要 · Abstract (English)
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation of English-French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that tests whether variation relationships are preserved across languages. Based on analyses of two parallel corpora, translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。