arXiv:2511.05162cs.CL2025-11被引 11

发现多语言模型性能差异实为翻译错误所致,修正后差距消失。

Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results

  • 分析多语言数学题集发现翻译错误导致性能假性差异
  • 提出半自动质检方法并改进答案提取流程,消除语言差距
  • 适合关注多语言模型公平性与评估可靠性的研究者

当前大型语言模型在多种语言(包括高资源与低资源语言)中表现优异,尤其在数学等跨领域任务上。本文以数学为例,研究不同语言下模型表现的差异。实验发现,模型在不同语言间存在显著且一致的性能差距,且该现象同时存在于高资源与低资源语言中。然而,进一步分析标准多语言数学基准(MGSM)后发现,数据中存在若干翻译错误,且模型输出的答案提取方式缺乏统一标准,共同扭曲了真实结果。为此,本文提出一种半自动质量保证方法以大规模修复翻译问题,并给出答案提取的改进建议。结合两项措施后,原本明显的语言差距基本消失,研究结论发生根本性转变。相关修正数据集已开源(https://github.com/google-research-datasets/MGSM-Rev2)。

原文摘要 · Abstract (English)

Most current large language models (LLMs) support a wide variety of languages in addition to English, including high-resource languages (e.g. German, Chinese, French), as well as low-resource ones (e.g. Swahili, Telugu). In addition, they have shown impressive capabilities in different domains, like coding, science and math. In this paper, taking math as an example domain, we study the performance of different LLMs across languages. Experimental results show that there exists a non-negligible and consistent gap in the performance of the models across languages. Interestingly, and somewhat against expectations, the gap exists for both high- and low-resource languages. These results should impact further research into cross-lingual capability generalization for next generation LLMs. Or they would, if it weren't for the fact that they are distorted by data quality issues. By analyzing one of the standard multilingual math benchmarks (MGSM), we determine that several translation errors are present in the data. Furthermore, the lack of standardized answer extraction from LLM outputs further influences the final results. We propose a method for semi-automatic quality assurance to address the first issue at scale, and give recommendations to address the second one. Combining these two approaches we show that the aforementioned language gap mostly disappears, leading to completely different conclusions from our research. We additionally release the corrected dataset to the community (https://github.com/google-research-datasets/MGSM-Rev2).

多语言模型评估偏差数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。