arXiv:2601.11778cs.CLcs.AI2026-01被引 6

用翻译质量当代理指标,低成本评估大模型多语言能力。

Translation as a Scalable Proxy for Multilingual Evaluation

  • 以翻译质量作为多语言能力的代理指标,降低评测成本。
  • 14个模型在9个基准上验证,翻译与下游任务相关性高达0.91。
  • 适合快速筛选多语言模型,尤其适用于资源稀缺语言。

大语言模型虽宣称具备多语言能力,但仅有不到30种语言拥有非机器翻译的完整基准测试,超过98%的世界语言(共7000种)仍处于实证空白。传统基准构建面临成本高、专家少、数据污染等问题。本文系统评估了翻译质量是否可作为模型多语言能力的可靠代理:对14个参数量从10亿到720亿的模型,在9个多样化基准和7种翻译评估指标上进行测试。结果显示,翻译表现与下游任务成功率高度相关(如Phi-4模型,中位皮尔逊相关系数MetricX=0.89,xCOMET=0.91,SSA-COMET=0.87)。这表明支持忠实翻译的表征能力,与多语言理解所需能力高度重叠。因此,翻译质量可作为低成本、高效的初步筛选代理,实现“先翻译、再聚焦”式评估。

原文摘要 · Abstract (English)

The rapid proliferation of LLMs has created a critical evaluation paradox: while LLMs claim multilingual proficiency, comprehensive non-machine-translated benchmarks exist for fewer than 30 languages, leaving >98% of the world's 7,000 languages in an empirical void. Traditional benchmark construction faces scaling challenges such as cost, scarcity of domain experts, and data contamination. We evaluate the validity of a simpler alternative: can translation quality alone indicate a model's broader multilingual capabilities? Through systematic evaluation of 14 models (1B-72B parameters) across 9 diverse benchmarks and 7 translation metrics, we find that translation performance is a good indicator of downstream task success (e.g., Phi-4, median Pearson r: MetricX = 0.89, xCOMET = 0.91, SSA-COMET = 0.87). These results suggest that the representational abilities supporting faithful translation overlap with those required for multilingual understanding. Translation quality, thus emerges as a strong, inexpensive first-pass proxy of multilingual performance, enabling a translation-first screening with targeted follow-up for specific tasks.

多语言评估翻译代理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。