测试大模型文化理解能力,发现语言流利不等于文化懂
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

- 11种语言、5个文化维度的原生问题数据集
- 模型文化能力依赖预训练暴露而非推理能力
- 现有推理优化方法无法真正解决文化偏差
多语言能力常隐含一个更强假设:能说用户语言的模型必然理解该语言所承载的文化。我们称之为‘文化对齐错觉’。为直接检验这一假设,我们提出MSQA,一个包含1,064个原生来源问题的多语言多文化简单问答基准,覆盖11个语言群体、5个文化维度和3个难度层级。不同于翻译生成的基准,MSQA聚焦本地化知识,减少英语中心的跨语言迁移捷径。评估18个大模型后发现,文化理解能力显著退化,且存在明显的局部效应:文化能力与预训练阶段的文化接触程度高度相关,远超一般推理能力的影响。进一步表明,常见的推理时修正手段无法消除此错觉:模型在不熟悉文化问题上仍过度自信,重复采样导致结果不稳定,检索增强对长尾事实帮助有限。这些发现表明,文化对齐不能仅从多语言能力推断,需要比校准、采样或检索更深层的干预。
原文摘要 · Abstract (English)
Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers. Unlike translated benchmarks, MSQA targets locally grounded knowledge and reduces shortcuts from English-centric cross-lingual transfer. Evaluating 18 LLMs, we find substantial cultural degradation and a pronounced Locality Effect: cultural competence tracks pre-training exposure more closely than general reasoning ability. We further show that common inference-time remedies do not dissolve the illusion. Models remain overconfident on unfamiliar cultural questions, repeated sampling yields unstable rather than reliable correctness, and retrieval augmentation helps unevenly on long-tail facts. These findings indicate that cultural alignment cannot be inferred from multilingual ability alone and requires deeper intervention than calibration, sampling, or retrieval at inference time
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。