arXiv:2510.05869cs.CL2025-10

质疑大模型文化倾向的稳定性,发现语言提示影响微乎其微

The fragility of "cultural tendencies" in LLMs

  • 用更广模型和更多题目复现实验,检验文化倾向是否真存在
  • 结果显示语言切换对输出影响极小,原结论不可靠
  • 适合关注AI文化偏见、方法论严谨性的研究者阅读

近期研究(Lu, Song, Zhang, 2025)声称,大语言模型在不同语言提示下表现出文化差异:中文提示时更偏向整体性与依存性,英文提示时更偏向分析性与独立性,并认为这是模型内嵌文化认知的表现。本文对该研究的方法、理论框架与结论提出质疑。我们通过扩大模型范围与测试项目数量进行针对性复现,结果表明提示语言对输出影响极为有限,所谓‘文化倾向’并非稳定特征,而是特定实验设计下的脆弱产物。研究指出,原结论过度解读了数据,模型并未编码真实文化信念。

原文摘要 · Abstract (English)

In a recent study, Lu, Song, and Zhang (2025) (LSZ) propose that large language models (LLMs), when prompted in different languages, display culturally specific tendencies. They report that the two models (i.e., GPT and ERNIE) respond in more interdependent and holistic ways when prompted in Chinese, and more independent and analytic ways when prompted in English. LSZ attribute these differences to deep-seated cultural patterns in the models, claiming that prompt language alone can induce substantial cultural shifts. While we acknowledge the empirical patterns they observed, we find their experiments, methods, and interpretations problematic. In this paper, we critically re-evaluate the methodology, theoretical framing, and conclusions of LSZ. We argue that the reported "cultural tendencies" are not stable traits but fragile artifacts of specific models and task design. To test this, we conducted targeted replications using a broader set of LLMs and a larger number of test items. Our results show that prompt language has minimal effect on outputs, challenging LSZ's claim that these models encode grounded cultural beliefs.

大模型文化偏见可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。