arXiv:2503.03308cs.CL2025-03被引 39

测试神经机器翻译的常识推理能力,发现模型表现普遍不佳。

The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation

  • 构建包含1200组对比翻译的常识测试集,覆盖7类常识类型。
  • 模型在三类歧义上的推理准确率仅60.1%,一致性31%。
  • 适合关注NMT常识缺陷的研究者和评测人员。

神经机器翻译是否符合常识?本文提出一个测试套件,用于评估神经机器翻译的常识推理能力。该套件包含三个测试集,涵盖词汇及句法歧义,需借助常识知识解决。我们手工创建了1200个三元组,每个包含源句和两组对比翻译,涉及7种常见常识类型。基于大规模语料预训练的语言模型(如BERT、GPT-2)在该测试套件的目标翻译上,常识推理准确率低于72%。我们在该套件上进行大量实验,评估神经机器翻译的常识推理能力,并分析影响因素。实验与分析表明,神经机器翻译在三类歧义上的推理准确率仅为60.1%,推理一致性低至31%。所构建的常识测试套件已开源:https://github.com/tjunlp-lab/CommonMT。

原文摘要 · Abstract (English)

Does neural machine translation yield translations that are congenial with common sense? In this paper, we present a test suite to evaluate the commonsense reasoning capability of neural machine translation. The test suite consists of three test sets, covering lexical and contextless/contextual syntactic ambiguity that requires commonsense knowledge to resolve. We manually create 1,200 triples, each of which contain a source sentence and two contrastive translations, involving 7 different common sense types. Language models pretrained on large-scale corpora, such as BERT, GPT-2, achieve a commonsense reasoning accuracy of lower than 72% on target translations of this test suite. We conduct extensive experiments on the test suite to evaluate commonsense reasoning in neural machine translation and investigate factors that have impact on this capability. Our experiments and analyses demonstrate that neural machine translation performs poorly on commonsense reasoning of the three ambiguity types in terms of both reasoning accuracy (60.1%) and reasoning consistency (31%). The built commonsense test suite is available at https://github.com/tjunlp-lab/CommonMT.

机器翻译常识推理评测数据集语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。