用维基数据自动评估大模型跨文化认知,发现英语下表现更好。
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
- 基于维基数据构建多语言自动评估框架,无需人工标注。
- 在13种语言、19个国家、6类主题中测试7个模型,发现英语下文化认知更强。
- 提出四维度指标,适合研究模型偏见与跨文化能力的研究者。
大型语言模型在全球多语言场景中广泛应用,但其以英语为主预训练导致跨语言文化认知差异,常产生偏见输出。现有评估受限于基准不足和翻译质量不可靠。为此,我们提出MAKIEval,一种基于维基数据的自动化多语言文化认知评估框架,可跨语言、跨地区、跨主题评估模型在开放文本生成中的文化知识表达。该框架利用维基数据的多语言结构作为跨语言锚点,自动识别模型输出中的文化实体并链接至结构化知识,实现无需人工标注或翻译的可扩展评估。我们设计了四项互补指标:粒度、多样性、文化特异性及跨语言一致性。评估涵盖7个来自不同地区的主流大模型(含开源与专有系统),覆盖13种语言、19个国家/地区及6个文化敏感主题(如食物、服饰)。结果显示,模型在英语提示下表现出更强的文化认知能力,表明英语更易激活文化相关知识。
原文摘要 · Abstract (English)
Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness, often resulting in biased outputs. However, comprehensive multilingual evaluation remains challenging due to limited benchmarks and questionable translation quality. To better assess these disparities, we introduce MAKIEval, an automatic multilingual framework for evaluating cultural awareness in LLMs across languages, regions, and topics. MAKIEval evaluates open-ended text generation, capturing how models express culturally grounded knowledge in natural language. Leveraging Wikidata's multilingual structure as a cross-lingual anchor, it automatically identifies cultural entities in model outputs and links them to structured knowledge, enabling scalable, language-agnostic evaluation without manual annotation or translation. We then introduce four metrics that capture complementary dimensions of cultural awareness: granularity, diversity, cultural specificity, and consensus across languages. We assess 7 LLMs developed from different parts of the world, encompassing both open-source and proprietary systems, across 13 languages, 19 countries and regions, and 6 culturally salient topics (e.g., food, clothing). Notably, we find that models tend to exhibit stronger cultural awareness in English, suggesting that English prompts more effectively activate culturally grounded knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。