arXiv:2510.14662cs.CL2025-10

教会机器翻译理解中文‘被’字句的语义倾向,提升负面内容翻译准确率。

Semantic Prosody in Machine Translation: the English-Chinese Case of Passive Structures

  • 构建中英句对数据集,展示‘被’字句的负面语义倾向。
  • 微调后模型在表达负面内容时更倾向使用‘被’字句,中性/正面则避免。
  • 多语言模型可跨语言迁移此语义知识,提升其他语种翻译表现。

语义倾向指语言单位与特定搭配词共现形成的语义色彩,应与字面意义区分。相同字面翻译的词语可能具有不同语义倾向,因此需关注此语言特性以生成准确翻译。然而现有机器翻译模型无法处理该问题。为此,我们提出一种方法,让机器翻译模型学习特定结构的语义倾向。聚焦中文‘被’字被动句,构建中英句子对数据集,以展示其负面语义倾向。随后,用该数据集微调 OPUS-MT、NLLB-600M 与 mBART50 模型。结果表明,微调后的模型在翻译负面内容时更倾向于使用‘被’字句,而在中性或正面内容中避免使用。此外,在多语言模型 NLLB-600M 中,该语义倾向知识可从英中翻译迁移至其他语种对,如西中翻译。

原文摘要 · Abstract (English)

Semantic prosody is a collocational meaning formed through the co-occurrence of a linguistic unit and a consistent series of collocates, which should be treated separately from semantic meaning. Since words that are literal translations of each other may have different semantic prosody, more attention should be paid to this linguistic property to generate accurate translations. However, current machine translation models cannot handle this problem. To bridge the gap, we propose an approach to teach machine translation models about semantic prosody of a specific structure. We focus on Chinese BEI passives and create a dataset of English-Chinese sentence pairs with the purpose of demonstrating the negative semantic prosody of BEI passives. Then we fine-tune OPUS-MT, NLLB-600M and mBART50 models with our dataset for the English-Chinese translation task. Our results show that fine-tuned MT models perform better on using BEI passives for translating unfavourable content and avoid using it for neutral and favourable content. Also, in NLLB-600M, which is a multilingual model, this knowledge of semantic prosody can be transferred from English-Chinese translation to other language pairs, such as Spanish-Chinese.

机器翻译语义倾向被动句跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。