arXiv:2602.20065cs.CLcs.AI2026-02被引 1

多语言大模型对不同语言的理解能力差异显著,英语并非最优。

Multilingual Large Language Models do not comprehend all natural languages to equal degrees

  • 在12种语言上测试主流模型的语义理解能力
  • 英语表现不如部分小语种,如罗曼语族语言
  • 模型在非英语语言上仍落后于人类,但差距各异

大型语言模型(LLMs)在信息获取中扮演关键角色,其核心能力依赖于对书面请求的理解。然而,当前对这一能力的认知受限于主要评估高资源语言(以西方、受教育、工业化、富裕、民主社区为主)的基准测试。普遍认为英语是性能最佳的语言,而小语种和低资源语言则被认为输出更不可靠,即便在最先进的多语言模型中也是如此。为追踪模型在不同语言中的理解能力差异,我们在12种语言上对3个主流模型进行了语言理解任务测试,涵盖印欧、亚非、突厥、汉藏和日语语系。结果显示,模型在类型多样语言中表现出显著的语言准确性,但在所有语言中均低于人类基线,且差距程度不一。出乎意料的是,英语并未表现最佳,反而被多个罗曼语族语言系统性超越,包括一些低资源语言。我们通过分析分词方式、与西班牙语和英语的语言距离、训练数据量、数据来源(高/低资源语言及WEIRD vs. 非WEIRD社区)等因素,解释了模型表现差异。

原文摘要 · Abstract (English)

Large Language Models (LLMs) play a critical role in how humans access information. While their core use relies on comprehending written requests, our understanding of this ability is currently limited, because most benchmarks evaluate LLMs in high-resource languages predominantly spoken by Western, Educated, Industrialised, Rich, and Democratic (WEIRD) communities. The default assumption is that English is the best-performing language for LLMs, while smaller, low-resource languages are linked to less reliable outputs, even in multilingual, state-of-the-art models. To track variation in the comprehension abilities of LLMs, we prompt 3 popular models on a language comprehension task across 12 languages, representing the Indo-European, Afro-Asiatic, Turkic, Sino-Tibetan, and Japonic language families. Our results suggest that the models exhibit remarkable linguistic accuracy across typologically diverse languages, yet they fall behind human baselines in all of them, albeit to different degrees. Contrary to what was expected, English is not the best-performing language, as it was systematically outperformed by several Romance languages, even lower-resource ones. We frame the results by discussing the role of several factors that drive LLM performance, such as tokenization, language distance from Spanish and English, size of training data, and data origin in high- vs. low-resource languages and WEIRD vs. non-WEIRD communities.

大模型多语言语言理解评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。