arXiv:2604.04732cs.CLcs.AI2026-04AAAI

检验大模型能否真正进行跨文化创作,而非仅做语言翻译。

Metaphors We Compute By: A Computational Audit of Cultural Translation vs. Thinking in LLMs

  • 用隐喻生成任务测试模型在五种文化中的创意表现
  • 发现模型对部分文化使用刻板隐喻,存在西方中心倾向
  • 提醒用户:指定文化身份不等于实现文化深度理解

大型语言模型(LLMs)常被称为多语言模型,因其能在多种语言中理解和回应。然而,会说一种语言并不等于能以该文化思维进行推理。这引出一个关键问题:大模型是否具备真正的文化意识推理能力?本文通过一项初步的计算审计,考察了在创造性写作任务中大模型的文化包容性。我们以跨越五个文化背景、针对多个抽象概念的隐喻生成任务为案例,实证检验模型是作为多元文化创作伙伴,还是仅作为依赖主导概念框架、仅做局部表达转换的文化翻译器。研究发现,模型在特定文化背景下表现出刻板化的隐喻使用,并存在西方默认倾向。这些结果表明,仅通过提示指定文化身份,并不能保证模型实现真正意义上的文化化推理。

原文摘要 · Abstract (English)

Large language models (LLMs) are often described as multilingual because they can understand and respond in many languages. However, speaking a language is not the same as reasoning within a culture. This distinction motivates a critical question: do LLMs truly conduct culture-aware reasoning? This paper presents a preliminary computational audit of cultural inclusivity in a creative writing task. We empirically examine whether LLMs act as culturally diverse creative partners or merely as cultural translators that leverage a dominant conceptual framework with localized expressions. Using a metaphor generation task spanning five cultural settings and several abstract concepts as a case study, we find that the model exhibits stereotyped metaphor usage for certain settings, as well as Western defaultism. These findings suggest that merely prompting an LLM with a cultural identity does not guarantee culturally grounded reasoning.

大模型文化偏见隐喻生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。