大模型语言能力不等于文化理解力,需额外努力才能避免美国中心偏见。
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
- 用世界价值观调查数据对比模型输出与各国真实民意
- 发现语言能力强的模型文化适配性未必好,尤其OpenAI模型
- 自我一致性比多语言能力更能预测文化契合度,适合做跨文化研究者参考
大型语言模型在多种语言上表现日益出色,但跨语言沟通能力并不等同于恰当的文化表达。一个关键问题是美国中心偏见——模型反映的是美国而非本地文化价值。本文提出一种新方法,将模型生成的回答分布与来自世界价值观调查(World Value Survey)的四国(丹麦、荷兰、英语、葡萄牙语)人群层面观点数据进行比较。采用严格的线性混合效应回归框架,对比了谷歌Gemma系列(2B–27B参数)和OpenAI Turbo系列多个迭代版本。结果显示,在不同模型家族间,语言能力与文化适配性之间无稳定关联。虽然Gemma模型在各语言中表现出语言能力与文化对齐的正相关,但OpenAI模型则无此趋势。更重要的是,自我一致性是文化多元适配性的更强预测因子,远超多语言能力本身。结果表明,实现真正意义上的文化对齐,需要超越提升通用语言能力的专门努力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are becoming increasingly capable across global languages. However, the ability to communicate across languages does not necessarily translate to appropriate cultural representations. A key concern is US-centric bias, where LLMs reflect US rather than local cultural values. We propose a novel methodology that compares LLM-generated response distributions against population-level opinion data from the World Value Survey across four languages (Danish, Dutch, English, and Portuguese). Using a rigorous linear mixed-effects regression framework, we compare two families of models: Google's Gemma models (2B--27B parameters) and successive iterations of OpenAI's turbo-series. Across the families of models, we find no consistent relationships between language capabilities and cultural alignment. While the Gemma models have a positive correlation between language capability and cultural alignment across languages, the OpenAI models do not. Importantly, we find that self-consistency is a stronger predictor of multicultural alignment than multilingual capabilities. Our results demonstrate that achieving meaningful cultural alignment requires dedicated effort beyond improving general language capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。