测试多语言模型对九国文化知识的掌握能力,发现英文表现反而优于当地语言。
Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
- 构建跨国家文化问答数据集XNationQA,覆盖7种语言、49,280个问题。
- 模型在英语中对非英语国家文化知识掌握更好,反向现象显著。
- 开源模型跨语言知识迁移能力极弱,适合评估文化公平性研究者使用。
现有多数多语言问答基准虽涵盖多种语言,但信息内容偏向西方,缺乏地域多样性,导致对多语言模型事实理解能力的评估存在偏差。为此,我们提出XNationQA,用于探究多语言大模型的文化素养。该数据集包含九个国家的地理、文化与历史相关问题共49,280个,以七种语言呈现。我们在八种主流多语言大模型上进行评测,并引入两种新型迁移评估指标。分析发现,模型在不同语言间获取特定文化事实的能力存在显著差异;尤其值得注意的是,模型在英语中对非英语国家文化的掌握程度常高于其母语。尽管模型在西方语言上表现更优,但这并不意味着其对西方国家更了解,这一结果出人意料。此外,模型跨语言知识迁移能力非常有限,尤其在开源模型中表现突出。
原文摘要 · Abstract (English)
Most multilingual question-answering benchmarks, while covering a diverse pool of languages, do not factor in regional diversity in the information they capture and tend to be Western-centric. This introduces a significant gap in fairly evaluating multilingual models' comprehension of factual information from diverse geographical locations. To address this, we introduce XNationQA for investigating the cultural literacy of multilingual LLMs. XNationQA encompasses a total of 49,280 questions on the geography, culture, and history of nine countries, presented in seven languages. We benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics. Our analyses uncover a considerable discrepancy in the models' accessibility to culturally specific facts across languages. Notably, we often find that a model demonstrates greater knowledge of cultural information in English than in the dominant language of the respective culture. The models exhibit better performance in Western languages, although this does not necessarily translate to being more literate for Western countries, which is counterintuitive. Furthermore, we observe that models have a very limited ability to transfer knowledge across languages, particularly evident in open-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。