arXiv:2503.07195cs.CL2025-03

用多源上下文提升翻译质量,发现特定场景下效果显著。

Contextual Cues in Machine Translation: Investigating the Potential of Multi-Source Input Strategies in LLMs and NMT Systems

  • 用中间语言翻译作上下文提示,增强目标翻译
  • 领域数据和远距离语言对中效果更好,高变异语料收益递减
  • 选高资源语言作上下文可提升性能,策略选型关键

我们研究了多源输入策略对机器翻译质量的影响,对比了GPT-4o(大语言模型)与传统多语言神经机器翻译(NMT)系统。通过使用中间语言翻译作为上下文线索,评估其在英译葡和中译葡任务中的有效性。结果表明,上下文信息显著提升了领域特定数据集及语言差异较大的语料翻译质量,但在语言变异性高的基准上出现收益递减。此外,我们在NMT系统中采用浅层融合的多源方法,发现以高资源语言作为上下文能有效提升其他翻译对的性能,凸显了上下文语言选择的战略重要性。

原文摘要 · Abstract (English)

We explore the impact of multi-source input strategies on machine translation (MT) quality, comparing GPT-4o, a large language model (LLM), with a traditional multilingual neural machine translation (NMT) system. Using intermediate language translations as contextual cues, we evaluate their effectiveness in enhancing English and Chinese translations into Portuguese. Results suggest that contextual information significantly improves translation quality for domain-specific datasets and potentially for linguistically distant language pairs, with diminishing returns observed in benchmarks with high linguistic variability. Additionally, we demonstrate that shallow fusion, a multi-source approach we apply within the NMT system, shows improved results when using high-resource languages as context for other translation pairs, highlighting the importance of strategic context language selection.

机器翻译多源输入上下文利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。