arXiv:2608.09766cs.CLcs.AI2026-08

构建本地化翻译评估集,发现模型对美国内容更优且易受数据污染。

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

  • 以语言对为单位的评估易污染,改用源语言对比和本地化设计。
  • 32个开源模型中,多数对非美地区内容表现更差,部分模型或过拟合训练集。
  • 适合关注多语种翻译鲁棒性与文化适配的研究者与开发者。

多语言翻译基准通常以英语为源语言并翻译成其他语言,将语言对作为评估单元,这种设计容易导致数据污染,并忽略地域与文化差异。为此,我们倡导源语言对比评估,并提出Cultivar——FLORES的一个本地化子集,支持针对特定地区的翻译评估。当与非本地化版本对比时,性能差异可用于探测数据污染和本地化鲁棒性。我们对32个开源模型进行了基准测试,发现专用机器翻译模型鲁棒性较弱,部分模型可能过拟合FLORES数据集,且无论语言如何,模型对美国内容的翻译表现普遍优于其他地区内容。

原文摘要 · Abstract (English)

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

机器翻译数据污染本地化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。