破解多语言评测中的文化偏见,提升全球公平性
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- 构建跨文化验证的多语言评测集,人工校验翻译质量
- 84.9%地理题聚焦北美欧洲,28%题目依赖西方文化知识
- 提供敏感与中立子集,适配跨国AI评估与模型对比
多语言数据集中的文化偏见严重影响其作为全球基准的有效性。这些偏见不仅源于语言差异,更来自对文化背景知识的需求,削弱了翻译数据集如MMLU的实际效用。机器翻译常引入语义扭曲或表达模糊。本文通过大规模评估主流开源与专有模型,发现MMLU上的进展高度依赖西方中心概念,28%的问题需文化敏感知识;涉及地理知识的问题中,84.9%聚焦北美洲或欧洲。模型排名随是否包含文化敏感题而变化,显示盲目使用翻译版MMLU会扭曲评估结果。为此,我们发布Global MMLU,覆盖42种语言,通过专业与社区标注员协同校验翻译质量,并系统评估原数据的文化偏见。该版本包含标注为文化敏感与文化中立的子集,支持更全面、公正的模型评估。
原文摘要 · Abstract (English)
Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also from the cultural knowledge required to interpret questions, reducing the practical utility of translated datasets like MMLU. Furthermore, translation often introduces artefacts that can distort the meaning or clarity of questions in the target language. A common practice in multilingual evaluation is to rely on machine-translated evaluation sets, but simply translating a dataset is insufficient to address these challenges. In this work, we trace the impact of both of these issues on multilingual evaluations and ensuing model performances. Our large-scale evaluation of state-of-the-art open and proprietary models illustrates that progress on MMLU depends heavily on learning Western-centric concepts, with 28% of all questions requiring culturally sensitive knowledge. Moreover, for questions requiring geographic knowledge, an astounding 84.9% focus on either North American or European regions. Rankings of model evaluations change depending on whether they are evaluated on the full portion or the subset of questions annotated as culturally sensitive, showing the distortion to model rankings when blindly relying on translated MMLU. We release Global MMLU, an improved MMLU with evaluation coverage across 42 languages -- with improved overall quality by engaging with compensated professional and community annotators to verify translation quality while also rigorously evaluating cultural biases present in the original dataset. This comprehensive Global MMLU set also includes designated subsets labeled as culturally sensitive and culturally agnostic to allow for more holistic, complete evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。