英语主导让本地文化知识被全球叙事覆盖,模型表现依赖语言环境。
When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models

- 构建717条孟加拉语问答数据集,含文化背景标注与双语对照。
- 英语提问使本地视角覆盖率下降30%以上,全球叙事占比显著上升。
- 引入本地证据可改善事实一致性,但无法根除语言带来的认知偏移。
大语言模型(LLMs)广泛用作跨语言知识接口,但带有文化根基的问题常反映全球主导叙事而非本地语境。我们以低资源文化语境的孟加拉语为例,研究这一失败模式——‘全球叙事主导’。提出 exttt{CulturalNB} 数据集,包含717条人工精标孟加拉文化实例,含平行孟加拉语—英语问答对、支持证据、元数据及社会文化标注。通过仅问题提示与基于证据提示,评估九种先进LLM,由人类与两名独立LLM评审员在跨语言一致性、语言锚定、全球替代、机构偏见及认识论视角覆盖等指标上进行评估。结果显示:英语提问系统性提升全球替代与机构化表述,降低本地视角覆盖率;本地证据可提升事实一致性和视角覆盖,但无法消除语言引发的认知转移。表明模型的文化缺陷不仅是知识缺失,更是语境锚定与叙事优先级的失效。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used as cross-lingual knowledge interfaces. However, culturally grounded questions often reflect globally dominant narratives rather than local contexts. We study this failure mode as \textit{global narrative dominance} in Bangla, a low-resource cultural context. We introduce \texttt{CulturalNB}, a dataset of 717 manually curated Bengali cultural instances with parallel Bangla--English question--answer pairs and supporting evidence, metadata, and sociocultural annotations. Using question-only and evidence-based prompting, we evaluate nine state-of-the-art LLMs with human and two independent LLM judges across metrics for cross-lingual consistency, language anchoring, global substitution, institutional bias, and epistemic perspective coverage. Results show that questions asked in English systematically increase global substitution and institutional framing while reducing local perspective coverage. Local evidence improves factual consistency and perspective coverage, but does not eliminate language-induced epistemic shifts. These findings suggest that cultural failures in LLMs are not only missing-knowledge errors but also failures of grounding and narrative prioritization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。