用多语言知识图谱提升跨文化实体翻译准确率
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
- 通过稠密检索融合多语言知识图谱信息,增强翻译模型理解力
- 在新构建的跨文化翻译数据集上,性能较NLLB-200提升129%
- 适合需要精准处理文化相关实体名称的翻译任务
包含实体名称的文本翻译极具挑战,因文化相关指涉在不同语言间差异显著,且常涉及超越音译与逐字翻译的再创作过程。本文从两方面应对跨文化翻译难题:(i) 构建首个大规模、人工标注的跨文化机器翻译基准数据集XC-Translate,聚焦具有文化敏感性的实体名称;(ii) 提出KG-MT,一种基于稠密检索机制,将多语言知识图谱信息端到端融入神经机器翻译模型的新方法。实验与分析表明,当前机器翻译系统及大语言模型在处理含实体名称文本时仍表现不佳,而KG-MT显著优于现有方法,相较NLLB-200实现129%相对提升,相较GPT-4提升62%。
原文摘要 · Abstract (English)
Translating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages. These variations may also be caused by transcreation, an adaptation process that entails more than transliteration and word-for-word translation. In this paper, we address the problem of cross-cultural translation on two fronts: (i) we introduce XC-Translate, the first large-scale, manually-created benchmark for machine translation that focuses on text that contains potentially culturally-nuanced entity names, and (ii) we propose KG-MT, a novel end-to-end method to integrate information from a multilingual knowledge graph into a neural machine translation model by leveraging a dense retrieval mechanism. Our experiments and analyses show that current machine translation systems and large language models still struggle to translate texts containing entity names, whereas KG-MT outperforms state-of-the-art approaches by a large margin, obtaining a 129% and 62% relative improvement compared to NLLB-200 and GPT-4, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。