剖析多语言大模型性能差异根源,揭示语言特征与模型表现的深层关联。
DEPART: DEcomposing PARiTy across Multilingual LLMs
- 构建贝叶斯分层模型,拆解多语言性能方差来源。
- 语言特征解释79%理解任务与92%推理任务的性能差异。
- 发现模型-评测交互是推理任务的主要差异驱动因素。
多语言大模型排行榜仅报告各语言准确率,却很少解释差异成因,导致系统性偏见无法归因,从业者也缺乏可操作改进方向。我们首先通过无分布假设的弗里德曼和克鲁斯卡尔-沃利斯检验,证明这些差距具有系统性而非采样噪声所致;随后提出两步贝叶斯分层框架,将多语言性能方差分解为可解释成分。第一步,分离语言身份带来的方差,发现语言特征(书写系统、语系、类型学距离)在理解任务中解释了79%的方差,在推理任务中解释了92%,且模型内部表示与英语的相似性成为两大任务的主导预测因子。第二步,对模型×基准×语言三维立方体进行分解,发现自然语言理解与推理任务的方差结构截然不同:模型身份主导理解任务(占66.7%方差),而基准×模型交互主导推理任务(占46.3%)。这些结果将多语言评估从被动性能映射转变为可解释、可诊断的框架,并提供精准干预路径。
原文摘要 · Abstract (English)
Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering practitioners no actionable levers. We first establish that these gaps are systematic rather than artifacts of sampling noise via distribution-free Friedman and Kruskal--Wallis tests, then introduce a two-step Bayesian hierarchical framework that decomposes multilingual performance variance into interpretable components. First, isolating the variance attributable to language identity, we show that observable language features (script, family, typological distance) explain $R^2_{\text{ling}} = 79\%$ of this variance on understanding tasks and $92\%$ on reasoning, with a model's internal representational similarity to English emerging as the dominant predictor across both task buckets. Second, decomposing the full (model$\times$benchmark$\times$language) cube, we find that NLU and reasoning have fundamentally divergent variance profiles: model identity dominates understanding ($66.7\%$ of variance), whereas the benchmark$\times$model interaction dominates reasoning ($46.3\%$). Together these results recast multilingual evaluation from passive performance mapping into an explainable, diagnostic framework with concrete levers for targeting the root drivers of language disparity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。