用人工校对数据优化术语翻译,让机器更懂企业语境。
Learning to Translate Ambiguous Terminology by Preference Optimization on Post-Edits
- 通过偏好优化学习校对数据中的术语选择逻辑。
- 术语准确率显著提升,且不影响整体翻译质量。
- 适合需要精准术语表达的企业级翻译场景。
实际翻译中,术语常无固定对应关系,同一术语在不同上下文可能有多个合法译法,正确性取决于企业风格指南和语境。神经机器翻译系统难以处理此类歧义。幸运的是,企业环境中存在大量关于有效但错误术语的校对数据。本文目标是基于这些校对记录学习术语消歧方法。提出一种偏好优化框架,以术语校对作为优先知识,无需依赖一对一术语词典或解码时的人工干预。在英德语后编辑数据上的实验表明,监督微调与偏好优化结合(含术语级和全句级目标),在不显著降低COMET得分的前提下,术语准确率相比强基线实现统计显著提升。此外,论文发布了测试集及术语词典。
原文摘要 · Abstract (English)
In real world translation scenarios, terminology is rarely one-to-one. Instead, multiple valid translations may appear in a terminology dictionary, but correctness of a translation depends on corporate style guides and context. This can be challenging for neural machine translation (NMT) systems. Luckily, in a corporate context, many examples of human post-edits of valid but incorrect terminology exist. The goal of this work is to learn how to disambiguate our terminology based on these corrections. Our approach is based on preference optimization, using the term post-edit as the knowledge to be preferred. While previous work had to rely on unambiguous translation dictionaries to set hard constraints during decoding, or to add soft constraints in the input, our framework requires neither one-to-one dictionaries nor human intervention at decoding time. We report results on English-German post-edited data and find that the optimal combination of supervised fine-tuning and preference optimization, with both term-specific and full sequence objectives, yields statistically significant improvements in term accuracy over a strong NMT baseline without significant losses in COMET score. Additionally, we release test sets from our post-edited data and terminology dictionary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。