提出新框架量化多语言大模型性能差异,更精准评估低资源语言表现。
Quantifying Language Disparities in Multilingual Large Language Models
- 分离实验变量,设计可解释的三项量化指标
- 发现高总性能不等于语言公平性高
- 特别适合评估低资源语言的模型表现
大规模多语言评估结果常因目标语言、实验设置和模型选择等因素而碎片化且混淆。本文提出一个解耦混杂变量的框架,引入三项可解释指标——性能实现率、变异系数和语言潜力,实现对模型与语言间实际性能差异的更细粒度、更深入量化。通过对13种模型变体在11个多语言数据集上的案例研究,验证该框架能更可靠地衡量模型性能与语言差异,尤其在低资源语言评估中表现突出。重要的是,结果表明更高整体性能并不意味着跨语言更公平。
原文摘要 · Abstract (English)
Results reported in large-scale multilingual evaluations are often fragmented and confounded by factors such as target languages, differences in experimental setups, and model choices. We propose a framework that disentangles these confounding variables and introduces three interpretable metrics--the performance realisation ratio, its coefficient of variation, and language potential--enabling a finer-grained and more insightful quantification of actual performance disparities across both (i) models and (ii) languages. Through a case study of 13 model variants on 11 multilingual datasets, we demonstrate that our framework provides a more reliable measurement of model performance and language disparities, particularly for low-resource languages, which have so far proven challenging to evaluate. Importantly, our results reveal that higher overall model performance does not necessarily imply greater fairness across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。