提升机器翻译自然度,兼顾内容准确与语言丰富性
Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation
- 多视角对齐机制增强翻译自然性
- 在荷兰语文学翻译中提升词汇丰富度且不损失准确率
- 适合需要高质量译文的评估数据构建场景
神经机器翻译系统会放大训练数据中的词汇偏见,导致输出语言过于贫乏。这类语言特征使自动翻译与原文及人工翻译差异明显,限制其在构建评估数据集等场景的应用。现有提升自然性的方法常以牺牲翻译准确性为代价。受人类反馈强化学习启发,本文提出一种新方法,同时奖励翻译自然性和内容保真度。通过多视角对齐生成更自然的译文,减少机器和人类翻译腔。在英译荷文学翻译任务上评估,最佳模型在保持翻译准确率的前提下,显著提升了词汇多样性与人类写作文本特征。
原文摘要 · Abstract (English)
Neural machine translation (NMT) systems amplify lexical biases present in their training data, leading to artificially impoverished language in output translations. These language-level characteristics render automatic translations different from text originally written in a language and human translations, which hinders their usefulness in for example creating evaluation datasets. Attempts to increase naturalness in NMT can fall short in terms of content preservation, where increased lexical diversity comes at the cost of translation accuracy. Inspired by the reinforcement learning from human feedback framework, we introduce a novel method that rewards both naturalness and content preservation. We experiment with multiple perspectives to produce more natural translations, aiming at reducing machine and human translationese. We evaluate our method on English-to-Dutch literary translation, and find that our best model produces translations that are lexically richer and exhibit more properties of human-written language, without loss in translation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。