arXiv:2608.06506cs.CL2026-08

发现语言模型在非英语问答中表现下降约17%,揭示了跨语言理解鸿沟。

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

论文配图:Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
图 1 · 摘自论文原文
  • 用平行语料库对比同内容不同语言下的模型表现,控制其他变量
  • 非英语回答的准确率比英语低17%(95%置信区间0.072-0.084)
  • 低资源语言模型表现更差,适合关注多语言公平性的研究者

语言模型常被评估为在英语中展现的能力在其他语言中同样存在。传统多语言基准很少在保持内容、问题、参考答案、模型和评估单元不变的前提下隔离语言因素。本文定义跨语言理解差距(CLCG)为相同内容与问题在目标语言中呈现时响应质量的下降。基于ParallelQA-18这一专业人工翻译的平行语料库,对五个实验室的五种模型在18种语言(英语为参考;葡萄牙语为高资源基线;16种目标语言覆盖Joshi等2020分类0-4类)的150篇文章中进行分层抽样评估。采用同项设计,仅改变篇章语言。主估计量对比英语与合并目标语言在高复杂度开放问答中的Token-F1微均值,使用文章聚类自助法计算置信区间。主要合并CLCG为0.078(95%置信区间0.072–0.084),相当于英语得分的约17%下降;同语言宏平均为0.077。剔除葡萄牙语后,宏平均差距为0.016(95%置信区间0.013–0.020)。语言级别的CLCG与Joshi资源类别负相关(rho = -0.594, p = 0.015, n = 16)。在盲测配对人类评估中,高资源语言回应在61.6%的决定性判断中更受青睐(估计偏好概率0.655,95%置信区间0.558–0.741)。英语中心的评估可能高估低资源语言用户的模型表现,模型能力不能简单从英语外推至其他语言。

原文摘要 · Abstract (English)

Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.

跨语言理解语言公平性评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。