arXiv:2508.20511cs.CLcs.AI2025-08EMNLP被引 6

现有多语言翻译评估基准存在文化偏见和质量漏洞,需更真实中立的测试标准。

Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark

  • 用自然语料替代领域特化文本,减少命名实体依赖。
  • 4种语言评估显示实际翻译质量低于90%标准。
  • 适合关注真实世界翻译挑战的研究者使用。

多语言机器翻译(MT)评估基准在评测现代MT系统能力中起核心作用。其中,FLORES+基准广泛使用,提供200多种语言的英译多语数据,并采用严格的质量控制。然而,我们对阿桑特蒂、日语、京帕沃和南阿塞拜疆语四门语言的研究揭示了该基准在真正多语言评估中的严重缺陷。人工评估显示,许多翻译未达宣称的90%质量标准,标注者指出源句常过于领域特定且文化偏向英语世界。我们进一步证明,仅通过复制命名实体等简单启发式方法即可获得非平凡的BLEU分数,表明评估协议存在漏洞。值得注意的是,基于高质量自然语料训练的模型在FLORES+上表现不佳,但在我们设计的领域相关评估集上取得显著提升。基于此,我们主张多语言评估应采用领域通用、文化中立的源文本,减少对命名实体的依赖,以更好反映真实翻译挑战。

原文摘要 · Abstract (English)

Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols. However, we study data in four languages (Asante Twi, Japanese, Jinghpaw, and South Azerbaijani) and uncover critical shortcomings in the benchmark's suitability for truly multilingual evaluation. Human assessments reveal that many translations fall below the claimed 90% quality standard, and the annotators report that source sentences are often too domain-specific and culturally biased toward the English-speaking world. We further demonstrate that simple heuristics, such as copying named entities, can yield non-trivial BLEU scores, suggesting vulnerabilities in the evaluation protocol. Notably, we show that MT models trained on high-quality, naturalistic data perform poorly on FLORES+ while achieving significant gains on our domain-relevant evaluation set. Based on these findings, we advocate for multilingual MT benchmarks that use domain-general and culturally neutral source texts rely less on named entities, in order to better reflect real-world translation challenges.

多语言翻译评估基准文化偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。