arXiv:2609.06634cs.CLcs.AI2026-09

用多语言语料暴露大模型翻译中词语搭配的盲区

Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus

论文配图:Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus
图 1 · 摘自论文原文
  • 基于AlphaMWE语料测试31个翻译系统表现
  • 意象化表达与固定搭配仍是翻译难点,人类评估更精准
  • 自动指标常失真,适合关注跨语言理解的读者

大语言模型在机器翻译任务中的表现往往依赖于其训练数据覆盖的领域和语种组合。为检验多词表达(MWEs)是否仍制约大模型的语言理解与翻译能力,我们基于公开的多语言平行语料库AlphaMWE,对WMT2026测试集共享任务中的31个翻译系统输出进行了评估。涵盖英译中(zh)、波兰语(pl)、德语(de)、阿拉伯语(ar)及两种方言(埃及、突尼斯阿拉伯语)。采用BLEU、ChrF、BERT-score进行自动评估,筛选出每语对前3名系统,并进一步开展人工评估。结果表明:意象化表达与多词搭配仍具挑战性;自动评估指标存在分歧;人工评估揭示了聚合分数掩盖的语言特异性错误。

原文摘要 · Abstract (English)

LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they are trained upon. To examine if Multiword Expressions (MWEs) still set a bottleneck for LLMs regarding language understanding and translation, we report the system performances from the WMT2026 Test Suites shared task, for which we used the publicly available multilingual parallel corpus AlphaMWE as the test suites. We received 31 MT systems' outputs covering English to Chinese (zh), Polish (pl), German (de), Arabic (ar) including Modern Standard Arabic (MSA) and two dialectal ones (Egyptian and Tunisian Arabic). We carried out automatic evaluations using BLEU, ChrF, BERT-score to select the Top3 systems per language pair, followed up with human evaluations on the selected systems. Our findings show that: figurative/MWE phenomena remain challenging; automatic metrics sometimes disagree; human evaluation uncovers language-specific errors hidden by aggregate scores.

机器翻译多词表达大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。