arXiv:2505.06010cs.CL2025-05中稿 · MTSummit 2025被引 1

测试主流翻译模型保留学名能力,发现表情符号等类别易出错

Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation -- a Multilingual Perspective

  • 构建多语言实体保留学名测试集,覆盖9类实体
  • 发现表情符号等特殊内容保留学名率不足60%
  • 适合研究跨语言实体保护的NLP工程师参考

当前机器翻译模型在多数场景下可生成高质量译文,但在保留特定实体(如网址、IBAN号、邮箱)方面仍存在问题。本文评估了OPUS、Google Translate、MADLAD和EuroLLM等主流NMT模型在英语、德语、波兰语和乌克兰语之间的实体保留学名能力。分析涵盖准确率、错误类型及成因,发现表情符号等类别对多数模型构成显著挑战。为此,我们构建了一个包含36,000句的多语言合成数据集,用于评估九类实体在四语言间的转移质量。

原文摘要 · Abstract (English)

Current machine translation models provide us with high-quality outputs in most scenarios. However, they still face some specific problems, such as detecting which entities should not be changed during translation. In this paper, we explore the abilities of popular NMT models, including models from the OPUS project, Google Translate, MADLAD, and EuroLLM, to preserve entities such as URL addresses, IBAN numbers, or emails when producing translations between four languages: English, German, Polish, and Ukrainian. We investigate the quality of popular NMT models in terms of accuracy, discuss errors made by the models, and examine the reasons for errors. Our analysis highlights specific categories, such as emojis, that pose significant challenges for many models considered. In addition to the analysis, we propose a new multilingual synthetic dataset of 36,000 sentences that can help assess the quality of entity transfer across nine categories and four aforementioned languages.

机器翻译实体保护多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。