arXiv:2604.03493cs.CL2026-04被引 1

用真人文化认知评估大模型输出,发现西方中心偏差和系统性错误。

Cultural Authenticity: Comparing LLM Cultural Representations to Native Human Expectations

  • 基于九国开放问卷构建文化重要性向量作为人类基准。
  • 三款前沿大模型在非美国家的文化对齐度随文化距离下降。
  • 发现所有模型存在高度相关(ρ>0.97)的系统性偏误。

大型语言模型(LLM)的文化表现通常通过文化多样性与事实准确性来评估,但缺乏对文化一致性的衡量:即生成内容是否反映本土人群对其文化要素的感知与优先级。本文提出一种以人类为中心的框架,评估模型输出与本地期望的一致性。首先,基于来自九个国家的开放式调查回复,构建了文化重要性向量(Cultural Importance Vectors)作为人类基准。其次,设计一种基于句法多样化提示集的方法,计算模型生成的文化表征向量(Cultural Representation Vectors),并应用于三款前沿模型(Gemini 2.5 Pro、GPT-4o、Claude 3.5 Haiku)。结果表明,部分模型呈现西方中心倾向,文化对齐度随国家与美国文化距离增加而降低。此外,所有模型均表现出高度相关的系统性误差签名(ρ > 0.97),过度强调某些文化符号,忽视用户深层社会价值与优先级。该方法推动评估从简单多样性转向真实捕捉全球文化复杂层级。

原文摘要 · Abstract (English)

Cultural representation in Large Language Model (LLM) outputs has primarily been evaluated through the proxies of cultural diversity and factual accuracy. However, a crucial gap remains in assessing cultural alignment: the degree to which generated content mirrors how native populations perceive and prioritize their own cultural facets. In this paper, we introduce a human-centered framework to evaluate the alignment of LLM generations with local expectations. First, we establish a human-derived ground-truth baseline of importance vectors, called Cultural Importance Vectors based on an induced set of culturally significant facets from open-ended survey responses collected across nine countries. Next, we introduce a method to compute model-derived Cultural Representation Vectors of an LLM based on a syntactically diversified prompt-set and apply it to three frontier LLMs (Gemini 2.5 Pro, GPT-4o, and Claude 3.5 Haiku). Our investigation of the alignment between the human-derived Cultural Importance and model-derived Cultural Representations reveals a Western-centric calibration for some of the models where alignment decreases as a country's cultural distance from the US increases. Furthermore, we identify highly correlated, systemic error signatures ($ρ> 0.97$) across all models, which over-index on some cultural markers while neglecting the deep-seated social and value-based priorities of users. Our approach moves beyond simple diversity metrics toward evaluating the fidelity of AI-generated content in authentically capturing the nuanced hierarchies of global cultures.

文化评估大模型偏差人类中心跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。